Skip to content

Preserve CSR backing during coordinate concatenation - #1005

Draft
MaykThewessen wants to merge 2 commits into
PyPSA:masterfrom
MaykThewessen:codex/csr-coordinate-concatenation
Draft

MaykThewessen wants to merge 2 commits into
PyPSA:masterfrom
MaykThewessen:codex/csr-coordinate-concatenation

Conversation

@MaykThewessen

Copy link
Copy Markdown
Contributor

Note

The following implementation summary and validation results were generated with AI assistance.

SnapshotWindow.merge concatenates expressions from disjoint snapshot blocks. Sparse inputs previously fell through to xarray concatenation, densifying their coefficients and mutating their backing. This change stacks eligible blocks in CSR form while preserving row order, coefficients, constants and coordinates.

The sparse path accepts ordinary expressions from the same model, all backed by CSR, with matching dimensions and an existing concatenation dimension containing unique, disjoint labels. Supported joins preserve auxiliary coordinates and absence semantics. Mixed inputs, repeated labels, unsupported expression classes and extra concatenation arguments retain the existing dense fallback.

Relates to #972.

Validation includes 18 new v1 cases. Focused tests passed 710 cases with 662 skips; Ruff passed. Local mypy was unavailable, so upstream CI remains necessary.

Benchmark and full-suite evidence

Three fresh-process repetitions concatenate four disjoint blocks of 25000 cells, with term widths [2, 2, 2, 128], using float64. Median process peak RSS falls from 0.571 to 0.438 GiB, a 23.3% reduction for this workload. Median merge time falls from 0.109 to 0.052 seconds. All output fingerprints match exactly. Timing and RSS are captured before fingerprinting.

A real BirdFlow 168-hour LP has 2.583.288 rows, 1.658.758 columns and 5.035.943 nonzeros. The matrix, objective coefficients, RHS, bounds, senses and variable/row labels match exactly. The densification warning falls from one to zero in each of three builds. Overall model performance is inconclusive because repeat ranges overlap. No independent solve or annual validation was performed.

Full-suite comparison against master f665a260 finds 94 baseline failures/errors versus 80 shared candidate failures/errors. The change resolves 14 regression-test failures and introduces none. The candidate still has 52 failures and 28 errors in this environment; the full suite is not green.

Implementation and this validation summary were written with GPT-6 assistance.

Concatenate constant and coordinate metadata with xarray, remap COO rows into the output grid, and retain explicit zeros. Keep overlapping labels and unsupported operands on the existing dense fallback.

Verify 18 targeted v1 cases and frozen constraints. Focused suite: 710 passed, 662 skipped. Full-suite comparison resolves 14 regression cases with no new failing test identities; 80 baseline failures/errors remain in the existing environment. Independently verify exact equality of a 168-hour BirdFlow LP and removal of its coordinate-merge densification warning. Ruff and read-only review pass.
@codspeed

codspeed Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Merging this PR will not alter performance

✅ 181 untouched benchmarks
⏩ 181 skipped benchmarks1


Comparing MaykThewessen:codex/csr-coordinate-concatenation (e3db8f4) with master (f665a26)

Open in CodSpeed

Footnotes

  1. 181 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

@MaykThewessen

Copy link
Copy Markdown
Contributor Author

Measured BirdFlow results for a pinned fork containing this work, compared with installed Linopy 0.9.1:

Measurement 0.9.1 dense/legacy Fork sparse/v1 + PyPSA repair
Whole-process peak RSS, sampled every 100 ms 6.1 GiB 5.7 GiB (-8%)
Representative solve wall time 1113 s 1058 s (-5%)
Completed windows / hours 6 / 1008 6 / 1008

The candidate was pinned to e3db8f403b664732a44ae9ad78598c601f1d08c0. Both arms used BirdFlow source 3faa673d, identical sealed inputs, float64, sequential 168-hour windows, HiGHS 1.15.1 HiPO with three threads and crossover, and disabled effective blueprint reuse. Source, input and solver distribution hashes stayed unchanged through both runs.

The complete 168-hour production LP matched exactly: sparse matrix plus objective, RHS, bounds, senses, variable/constraint labels and variable types (all continuous for production linearized UC). This includes 492 minimum-up/down-time rows restored by an isolated PyPSA producer repair. Shifted pre-window terms became absent under v1 semantics, which dropped nonempty UC rows. The repair makes those boundary expression terms mathematical zero. Its 46 tests passed against the fork; all 40 existing UC tests also passed with v1 globally enabled.

Across all 1008 hours, dispatch, storage state, flows, nodal prices and operative loading matched exactly, with zero slack and unchanged congestion counts. All six objective values and actual cached output file hashes matched; exports and representative health checks passed.

Limits: this compares the entire pinned fork plus the PyPSA repair, so it does not isolate PR #1005's gain. Timing is one sequential A/B pair. Both retain the same existing 1.24 kW rating excess and DC-angle warning. No annual or worker-scaling claim, and these results do not establish compatibility of later packaging changes.

Note: AI-assisted (written with GPT-6).

@FabianHofmann

Copy link
Copy Markdown
Collaborator

thanks @MaykThewessen! I already had this on the radar and tackled it in #1016 ; closing this now

@MaykThewessen

Copy link
Copy Markdown
Contributor Author

ah sorry didn't check for existing issues

Great that you tackled it in #1016 now!

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants