Skip to content

Preserve weights across pandas operations and correct estimator docs - #336

Open
juaristi22 wants to merge 2 commits into
mainfrom
fix-weighted-operations-333-335
Open

juaristi22 wants to merge 2 commits into
mainfrom
fix-weighted-operations-333-335

Conversation

@juaristi22

@juaristi22 juaristi22 commented Sep 22, 2026 •

Copy link
Copy Markdown
Collaborator

Ordinary pandas spellings such as groupby(...).agg({"x": "mean"}), callable pivots, and row-wise apply could drop observation weights. On values [10, 100] with weights [9, 1], these paths now return the weighted mean of 19. Unsupported estimators raise with an explicit plain-pandas escape, and binary operations reject conflicting weights.

The change also preserves source weights in constructors and pandas conversions, fixes numeric-only reductions and numeric-frame dropna, and adds weighted value_counts and mode. The regression cases generate a support matrix linked from the README and docs. CI explicitly tests pandas 2 and 3.

Poverty-method docstrings now identify weighted headcount rates, aggregate currency gaps, and aggregate squared-currency gaps. Normalised FGT(1) and FGT(2) remain outside the API's scope. Hand-calculated tests cover all five estimators with non-uniform weights, a threshold-boundary row, and a zero-weight row. The covariance example now describes frequency weighting and verifies its numbers against NumPy.

Remaining #333 boundary: when a plain pandas object comes first in pd.concat, pandas selects its constructor without calling microdf's subclass hooks. This PR adds microdf.concat, which validates every input before pandas dispatches, and records the direct-pandas limitation in the support matrix and a strict expected-failure regression. It does not monkey-patch pandas globally, and intentionally leaves #333 open.

Validation: 991 tests passed and one expected failure on each of pandas 2.3.3 and 3.0.6 (Python 3.13.14). The same 991 tests passed on Python 3.9.6 with NumPy 1.26.4 and pandas 2.3.3. make format, make lint, git diff --check, and the MyST HTML documentation build passed. The expected failure is solely the direct plain-first pandas concat case described above.

Addresses #333.
Fixes #334.
Fixes #335.

Credit to @baogorek for the original aggregation and weight-loss reports in #264 and #265.

@juaristi22
juaristi22 marked this pull request as ready for review September 22, 2026 15:34
@MaxGhenis

Copy link
Copy Markdown
Collaborator

Please hold off merging this for now: it breaks PolicyEngine simulations. On this branch np.all, np.where, np.select, np.unique and np.bincount raise NotImplementedError on a MicroSeries; on 1.5.10 they work. policyengine-core calls np.all(~mask) on a MicroSeries for every variable with defined_for (policyengine_core/simulations/simulation.py:822), and 15 policyengine.py tests that pass on 1.5.10 fail on this branch. policyengine-core and policyengine-us accept any microdf from 1.0, policyengine-uk any from 1.2.1, and a merge to main publishes to PyPI, so this would reach new installs straight away. Full review to follow.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants