Preserve weights across pandas operations and correct estimator docs - #336
Open
juaristi22 wants to merge 2 commits into
Open
juaristi22 wants to merge 2 commits into
juaristi22 wants to merge 2 commits into
Conversation
juaristi22
marked this pull request as ready for review
September 22, 2026 15:34
Collaborator
|
Please hold off merging this for now: it breaks PolicyEngine simulations. On this branch |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Ordinary pandas spellings such as
groupby(...).agg({"x": "mean"}), callable pivots, and row-wiseapplycould drop observation weights. On values[10, 100]with weights[9, 1], these paths now return the weighted mean of 19. Unsupported estimators raise with an explicit plain-pandas escape, and binary operations reject conflicting weights.The change also preserves source weights in constructors and pandas conversions, fixes numeric-only reductions and numeric-frame
dropna, and adds weightedvalue_countsandmode. The regression cases generate a support matrix linked from the README and docs. CI explicitly tests pandas 2 and 3.Poverty-method docstrings now identify weighted headcount rates, aggregate currency gaps, and aggregate squared-currency gaps. Normalised FGT(1) and FGT(2) remain outside the API's scope. Hand-calculated tests cover all five estimators with non-uniform weights, a threshold-boundary row, and a zero-weight row. The covariance example now describes frequency weighting and verifies its numbers against NumPy.
Remaining #333 boundary: when a plain pandas object comes first in
pd.concat, pandas selects its constructor without calling microdf's subclass hooks. This PR addsmicrodf.concat, which validates every input before pandas dispatches, and records the direct-pandas limitation in the support matrix and a strict expected-failure regression. It does not monkey-patch pandas globally, and intentionally leaves #333 open.Validation: 991 tests passed and one expected failure on each of pandas 2.3.3 and 3.0.6 (Python 3.13.14). The same 991 tests passed on Python 3.9.6 with NumPy 1.26.4 and pandas 2.3.3.
make format,make lint,git diff --check, and the MyST HTML documentation build passed. The expected failure is solely the direct plain-first pandas concat case described above.Addresses #333.
Fixes #334.
Fixes #335.
Credit to @baogorek for the original aggregation and weight-loss reports in #264 and #265.