From 66beb3d53bc3401cda438d03d5c01c8b4b6076b7 Mon Sep 17 00:00:00 2001 From: Vahid Ahmadi Date: Mon, 14 Sep 2026 18:16:09 +0100 Subject: [PATCH 01/22] Add a JOSS paper and the repository files a submission needs The paper is framed around the two problems microdf solves: the estimator decisions that hand-written weighting makes implicitly, and keeping weights aligned with the rows they describe through merges, filters and grouping, where a misalignment raises nothing and leaves a plausible wrong answer. Also adds a citation file, a code of conduct, and the workflow that builds a draft PDF, all of which a JOSS review checks for. The author list and ORCIDs still need confirming before submission. Co-Authored-By: Claude Opus 5 (1M context) --- .github/workflows/draft-pdf.yml | 23 +++++++ CITATION.cff | 27 ++++++++ CODE_OF_CONDUCT.md | 64 +++++++++++++++++++ changelog.d/joss-paper.added.md | 1 + paper.bib | 109 ++++++++++++++++++++++++++++++++ paper.md | 102 ++++++++++++++++++++++++++++++ 6 files changed, 326 insertions(+) create mode 100644 .github/workflows/draft-pdf.yml create mode 100644 CITATION.cff create mode 100644 CODE_OF_CONDUCT.md create mode 100644 changelog.d/joss-paper.added.md create mode 100644 paper.bib create mode 100644 paper.md diff --git a/.github/workflows/draft-pdf.yml b/.github/workflows/draft-pdf.yml new file mode 100644 index 00000000..d057c09a --- /dev/null +++ b/.github/workflows/draft-pdf.yml @@ -0,0 +1,23 @@ +on: + push: + paths: + - paper.md + - paper.bib + pull_request: + paths: + - paper.md + - paper.bib + +jobs: + paper: + runs-on: ubuntu-latest + name: Draft PDF + steps: + - uses: actions/checkout@v4 + - uses: openjournals/openjournals-draft-action@master + with: + journal: joss + - uses: actions/upload-artifact@v4 + with: + name: paper + path: paper.pdf diff --git a/CITATION.cff b/CITATION.cff new file mode 100644 index 00000000..6d3660cf --- /dev/null +++ b/CITATION.cff @@ -0,0 +1,27 @@ +cff-version: 1.2.0 +message: "If you use this software, please cite it as below." +type: software +title: "microdf: Weighted DataFrames and Series for survey microdata analysis" +url: "https://github.com/PolicyEngine/microdf" +repository-code: "https://github.com/PolicyEngine/microdf" +abstract: "Weighted pandas DataFrames and Series for survey microdata, with distributional statistics including weighted quantiles, Gini, income shares and Foster-Greer-Thorbecke poverty measures." +license: MIT +authors: + - family-names: Ahmadi + given-names: Vahid + orcid: "https://orcid.org/0009-0004-1093-6272" + affiliation: "PolicyEngine, Washington, DC, United States" + - family-names: Ghenis + given-names: Max + orcid: "https://orcid.org/0000-0002-1335-8277" + affiliation: "PolicyEngine, Washington, DC, United States" + - family-names: Juaristi + given-names: María + affiliation: "PolicyEngine, Washington, DC, United States" +keywords: + - survey statistics + - microdata + - inequality + - poverty + - pandas + - Python diff --git a/CODE_OF_CONDUCT.md b/CODE_OF_CONDUCT.md new file mode 100644 index 00000000..56cdab0f --- /dev/null +++ b/CODE_OF_CONDUCT.md @@ -0,0 +1,64 @@ +# Contributor covenant code of conduct + +## Our pledge + +We as members, contributors, and leaders pledge to make participation in our +community a harassment-free experience for everyone, regardless of age, body +size, visible or invisible disability, ethnicity, sex characteristics, gender +identity and expression, level of experience, education, socio-economic status, +nationality, personal appearance, race, religion, or sexual identity +and orientation. + +We pledge to act and interact in ways that contribute to an open, welcoming, +diverse, inclusive, and healthy community. + +## Our standards + +Examples of behavior that contributes to a positive environment for our +community include: + +* Demonstrating empathy and kindness toward other people +* Being respectful of differing opinions, viewpoints, and experiences +* Giving and gracefully accepting constructive feedback +* Accepting responsibility and apologizing to those affected by our mistakes, + and learning from the experience +* Focusing on what is best not just for us as individuals, but for the + overall community + +Examples of unacceptable behavior include: + +* The use of sexualized language or imagery, and sexual attention or + advances of any kind +* Trolling, insulting or derogatory comments, and personal or political attacks +* Public or private harassment +* Publishing others' private information, such as a physical or email + address, without their explicit permission +* Other conduct which could reasonably be considered inappropriate in a + professional setting + +## Enforcement responsibilities + +Community leaders are responsible for clarifying and enforcing our standards of +acceptable behavior and will take appropriate and fair corrective action in +response to any behavior that they deem inappropriate, threatening, offensive, +or harmful. + +## Scope + +This Code of Conduct applies within all community spaces, and also applies when +an individual is officially representing the community in public spaces. + +## Enforcement + +Instances of abusive, harassing, or otherwise unacceptable behavior may be +reported to the community leaders responsible for enforcement at +hello@policyengine.org. All complaints will be reviewed and investigated +promptly and fairly. + +## Attribution + +This Code of Conduct is adapted from the [Contributor Covenant][homepage], +version 2.0, available at +https://www.contributor-covenant.org/version/2/0/code_of_conduct.html. + +[homepage]: https://www.contributor-covenant.org diff --git a/changelog.d/joss-paper.added.md b/changelog.d/joss-paper.added.md new file mode 100644 index 00000000..abbe4e56 --- /dev/null +++ b/changelog.d/joss-paper.added.md @@ -0,0 +1 @@ +- Added a JOSS paper, a citation file, a code of conduct, and a workflow that builds a draft PDF of the paper. diff --git a/paper.bib b/paper.bib new file mode 100644 index 00000000..c6e092f1 --- /dev/null +++ b/paper.bib @@ -0,0 +1,109 @@ +@software{policyengine_py, + title={{policyengine}}, + author={{PolicyEngine Contributors}}, + version={4.10.0}, + url={https://github.com/PolicyEngine/policyengine.py}, + year={2026} +} + +@misc{arnold_ventures, + title={Public Finance Program}, + author={{Arnold Ventures}}, + year={2023}, + note={Grant to PolicyEngine for congressional district-level policy analysis}, + url={https://www.arnoldventures.org/work/public-finance} +} + +@misc{nsf_pose, + title={{POSE}: Phase {I}: {PolicyEngine} -- Advancing Public Policy Analysis}, + author={{National Science Foundation}}, + year={2025}, + note={Award 2518372. PI: Max Ghenis, PSL Foundation. \$299,974}, + url={https://www.nsf.gov/awardsearch/showAward.jsp?AWD_ID=2518372} +} + +@online{neo_philanthropy, + title={{NEO Philanthropy} Awards \$200,000 Grant to {PolicyEngine}}, + author={Ghenis, Max}, + year={2024}, + url={https://policyengine.org/us/research/neo-philanthropy} +} + +@software{claude2026, + title={{Claude}}, + author={{Anthropic}}, + year={2026}, + note={Opus 4 model used for code refactoring assistance}, + url={https://www.anthropic.com/claude} +} + +@misc{nuffield2024grant, + title={Enhancing, localising and democratising tax-benefit policy analysis}, + author={{Nuffield Foundation}}, + year={2024}, + note={General Election Analysis and Briefing Fund grant to PolicyEngine}, + url={https://www.nuffieldfoundation.org/project/enhancing-localising-and-democratising-tax-benefit-policy-analysis} +} +@inproceedings{mckinney2010pandas, + author = {McKinney, Wes}, + title = {Data Structures for Statistical Computing in Python}, + booktitle = {Proceedings of the 9th Python in Science Conference}, + pages = {56--61}, + year = {2010}, + doi = {10.25080/Majora-92bf1922-00a} +} + +@software{pandas2020, + author = {{The pandas development team}}, + title = {pandas-dev/pandas: Pandas}, + publisher = {Zenodo}, + year = {2020}, + doi = {10.5281/zenodo.3509134}, + url = {https://doi.org/10.5281/zenodo.3509134} +} + +@inproceedings{seabold2010statsmodels, + author = {Seabold, Skipper and Perktold, Josef}, + title = {statsmodels: Econometric and statistical modeling with Python}, + booktitle = {Proceedings of the 9th Python in Science Conference}, + year = {2010}, + doi = {10.25080/Majora-92bf1922-011} +} + +@article{foster1984fgt, + author = {Foster, James and Greer, Joel and Thorbecke, Erik}, + title = {A Class of Decomposable Poverty Measures}, + journal = {Econometrica}, + volume = {52}, + number = {3}, + pages = {761--766}, + year = {1984}, + doi = {10.2307/1913475} +} + +@article{lumley2004survey, + author = {Lumley, Thomas}, + title = {Analysis of Complex Survey Samples}, + journal = {Journal of Statistical Software}, + volume = {9}, + number = {8}, + pages = {1--19}, + year = {2004}, + doi = {10.18637/jss.v009.i08} +} + +@book{lumley2010complex, + author = {Lumley, Thomas}, + title = {Complex Surveys: A Guide to Analysis Using R}, + publisher = {John Wiley and Sons}, + year = {2010}, + doi = {10.1002/9780470580066} +} + +@software{samplics, + author = {Diallo, Mamadou S.}, + title = {samplics: Sampling techniques for complex survey designs}, + year = {2021}, + url = {https://github.com/samplics-org/samplics}, + doi = {10.21105/joss.03376} +} diff --git a/paper.md b/paper.md new file mode 100644 index 00000000..27d4c70e --- /dev/null +++ b/paper.md @@ -0,0 +1,102 @@ +--- +title: "microdf: Weighted DataFrames and Series for survey microdata analysis" +tags: + - Python + - survey statistics + - microdata + - inequality + - poverty + - pandas +authors: + - name: Vahid Ahmadi + orcid: 0009-0004-1093-6272 + affiliation: '1' + corresponding: true + - name: Max Ghenis + orcid: 0000-0002-1335-8277 + affiliation: '1' + - name: María Juaristi + affiliation: '1' +affiliations: + - name: PolicyEngine, Washington, DC, United States + index: '1' +date: 14 September 2026 +bibliography: paper.bib +--- + +# Summary + +`microdf` provides weighted data structures for survey microdata analysis in Python. Survey records carry sampling weights: each row stands for many households, and the weights vary by orders of magnitude within a single file. Statistics computed without them describe the sample rather than the population, and the two can differ substantially. + +The package's central design choice is that the weight is a property of the data structure rather than an argument to a function. `MicroSeries` and `MicroDataFrame` subclass the pandas [@mckinney2010pandas; @pandas2020] structures and carry a weight vector through the operations an analysis pipeline performs — indexing, merging, grouping, reindexing, type conversion — so that a statistic computed at the end of a pipeline is weighted whether or not the analyst remembers to weight it. Aggregations that pandas defines but that would silently ignore weights are overridden rather than inherited, so a method either returns a weighted result or warns that it cannot. + +On top of that foundation the package implements the estimators distributional analysis reports: quantiles by the inverse cumulative distribution function, frequency-weighted variance, the Gini coefficient from the Lorenz curve, top and bottom shares with proportional handling of records tied at the cutoff, and the Foster-Greer-Thorbecke family of poverty measures [@foster1984fgt]. + +# Statement of Need + +Analysts working with survey microdata in Python face two distinct problems, and the second is the one that causes silent errors. + +The first is that the estimators themselves require care. A weighted median is not the median of the weighted values. A weighted variance requires deciding whether weights are frequencies or precision weights, and the two give different answers. A top-1% share requires deciding what happens to the records straddling the cutoff: assigning them wholly to one side introduces a bias that grows as weights grow coarser. These are solvable problems, but each is a decision that hand-written code makes implicitly, usually without recording it, so implementations diverge between analysts on exactly the edge cases that matter. + +The second problem is that weights must stay aligned with the data through every transformation before the estimator runs. Building an analysis dataset means merging administrative variables onto survey records, filtering to a subpopulation, grouping by geography, reindexing after a sort. Each of these is an opportunity for the weight vector to fall out of alignment with the rows it describes, and nothing raises when it does: the pipeline completes and returns a number that is wrong by an amount nobody can see. In our experience maintaining microsimulation datasets, this is a more frequent source of error than the estimator formulas, and a harder one to detect, because the result remains plausible. + +`microdf` addresses both. It makes the estimator decisions once, documents them, and tests them: the quantile estimator follows the inverse CDF definition, matching the default behaviour of R's `survey::svyquantile` [@lumley2004survey; @lumley2010complex] so that results can be checked against an established implementation; the variance treats weights as frequencies, so that with integer weights it agrees with `numpy` on the replicated sample; the top-share estimator splits the record at the cutoff proportionally, so a constant distribution returns the share it should. And it keeps the weight attached to the data across the transformations between loading a file and computing a statistic, which is what makes those estimator guarantees worth anything in a real pipeline. + +# State of the Field + +Several tools compute weighted statistics, but the combination `microdf` occupies — pandas-native structures, a distributional estimator set, and no complex-survey design object — is not otherwise filled. + +| Tool | Weighted quantiles | Inequality and poverty measures | pandas-native | Design-based variance | +|---|---|---|---|---| +| `microdf` | Yes | Gini, top and bottom shares, FGT poverty | Yes | No | +| `samplics` [@samplics] | Yes | No | Partly | Yes | +| `statsmodels` [@seabold2010statsmodels] | Limited | No | Partly | Partly | +| R `survey` [@lumley2004survey] | Yes | Limited | No (R) | Yes | +| pandas + manual weighting | Hand-written | Hand-written | Yes | No | + +R's `survey` package is the reference implementation for design-based survey inference and remains the right tool when standard errors under a complex design are required. `samplics` brings much of that machinery to Python, again centred on sampling design. Neither is built around the inequality and poverty estimators that distributional policy analysis reports, and neither returns objects that behave like a `DataFrame` in an existing pandas pipeline. + +`microdf` deliberately does not implement design-based variance estimation. Its weights are population weights: it estimates the statistic, not the sampling error around it. Analysts needing standard errors under a stratified or clustered design should use `survey` or `samplics`. The trade is a much smaller interface, and statistics that compose with the pandas code analysts already have. + +# Software Design + +`MicroSeries` extends `pandas.Series` with a weight vector of equal length; `MicroDataFrame` extends `pandas.DataFrame`, holds a weight column, and exposes each column as a `MicroSeries`. + +Pandas methods are classified into three groups, and the classification is what makes the guarantee hold. *Scalar* methods return a single weighted statistic and are overridden to use the weights. *Vector* methods return a series aligned to the input and carry the weights through to the result. *Agnostic* methods do not depend on weighting and are inherited unchanged. Methods that would need weighting but do not yet implement it, currently `cov` and `corr`, fall through to pandas and emit a warning rather than returning an unweighted number silently. Shape-changing operations — `merge`, `groupby`, `reset_index`, `drop`, `astype` — are overridden so the weight vector follows the rows it describes. + +```python +import microdf as mdf + +df = mdf.MicroDataFrame( + {"income": [10_000, 30_000, 120_000], "threshold": [15_000, 15_000, 15_000]}, + weights=[800, 1_200, 50], +) + +df.income.median() # 30000.0, weighted +df.income.gini() # Lorenz-curve Gini over the weighted distribution +df.income.top_10_pct_share() # proportional split at the cutoff +df.poverty_rate("income", "threshold") + +regional = df.merge(geography, on="household_id") # weights follow the join +regional.groupby("region").income.median() # and the grouping +``` + +Statistics requiring a decision take it as an explicit argument rather than choosing silently: `gini` accepts a `negatives` policy, `var` takes `ddof`, and behaviour on zero-weight records is documented and tested. The estimators are small and independently checkable: weighted quantiles sort by value, accumulate weight, and return the smallest value whose cumulative weight share reaches `q`, after dropping zero-weight records so they cannot be selected; the Gini is computed from the Lorenz curve over weighted cumulative population and income; poverty measures follow the FGT family, with rate, gap, deep gap, and squared gap. + +# Research Impact Statement + +`microdf` has been public since June 2018, with 749 commits across 16 contributors and 11 releases. It is a dependency of both `policyengine-us` and `policyengine-uk`, and therefore sits in the computational path of PolicyEngine's published distributional estimates — the poverty rates, decile impacts, and Gini changes reported in its analyses and through its web application [@policyengine_py]. It is also used directly in standalone policy studies, including analyses of free school meals, extended childcare entitlements, and national insurance reforms. + +The package's role is that of infrastructure: it is not the visible output of an analysis, but the layer that determines whether a reported poverty rate is a population estimate or a sample artefact. Its adoption is best measured by the analyses that depend on it rather than by direct use. + +# Acknowledgements + +Arnold Ventures [@arnold_ventures], NEO Philanthropy [@neo_philanthropy], the Gerald Huff Fund for Humanity, and the National Science Foundation (NSF POSE Phase I, Award 2518372) [@nsf_pose] funded this work in the US. The Nuffield Foundation has funded the UK work since September 2024 [@nuffield2024grant]. These funders had no involvement in the design, development, or content of this software or paper. All authors are employed by PolicyEngine and may benefit reputationally from the software's adoption; this relationship is disclosed here as a potential conflict of interest. + +We thank all contributors to `microdf`, and Thomas Lumley, whose `survey` package provides the reference against which several of the estimators here are checked. + +# AI Usage Disclosure + +The authors used generative AI tools, specifically Claude Opus by Anthropic [@claude2026], to assist with code refactoring, test authoring, and drafting of this paper. Human authors reviewed, edited, and validated all AI-assisted outputs, and made all decisions regarding estimator definitions and software design. The authors remain fully responsible for the accuracy, originality, and correctness of all submitted materials. + +# References From 600f57b69f90360faac32fa384dfa5a10a478fce Mon Sep 17 00:00:00 2001 From: Vahid Ahmadi Date: Wed, 16 Sep 2026 10:40:15 +0100 Subject: [PATCH 02/22] Use title case, and describe weight preservation as it actually works The draft said a statistic is weighted whether or not the analyst remembers to weight it. Issue #300 records that weight survival is guaranteed only for the operations explicitly overridden, so the paper now says that, names them, and notes that extending the set is ongoing work. A reviewer reads the issue tracker, and a claim the tracker contradicts is worse than a narrower one. Co-Authored-By: Claude Opus 5 (1M context) --- CITATION.cff | 2 +- paper.md | 10 ++++++---- 2 files changed, 7 insertions(+), 5 deletions(-) diff --git a/CITATION.cff b/CITATION.cff index 6d3660cf..50e6c4b0 100644 --- a/CITATION.cff +++ b/CITATION.cff @@ -1,7 +1,7 @@ cff-version: 1.2.0 message: "If you use this software, please cite it as below." type: software -title: "microdf: Weighted DataFrames and Series for survey microdata analysis" +title: "microdf: Weighted DataFrames and Series for Survey Microdata Analysis" url: "https://github.com/PolicyEngine/microdf" repository-code: "https://github.com/PolicyEngine/microdf" abstract: "Weighted pandas DataFrames and Series for survey microdata, with distributional statistics including weighted quantiles, Gini, income shares and Foster-Greer-Thorbecke poverty measures." diff --git a/paper.md b/paper.md index 27d4c70e..71d14551 100644 --- a/paper.md +++ b/paper.md @@ -1,5 +1,5 @@ --- -title: "microdf: Weighted DataFrames and Series for survey microdata analysis" +title: "microdf: Weighted DataFrames and Series for Survey Microdata Analysis" tags: - Python - survey statistics @@ -28,7 +28,7 @@ bibliography: paper.bib `microdf` provides weighted data structures for survey microdata analysis in Python. Survey records carry sampling weights: each row stands for many households, and the weights vary by orders of magnitude within a single file. Statistics computed without them describe the sample rather than the population, and the two can differ substantially. -The package's central design choice is that the weight is a property of the data structure rather than an argument to a function. `MicroSeries` and `MicroDataFrame` subclass the pandas [@mckinney2010pandas; @pandas2020] structures and carry a weight vector through the operations an analysis pipeline performs — indexing, merging, grouping, reindexing, type conversion — so that a statistic computed at the end of a pipeline is weighted whether or not the analyst remembers to weight it. Aggregations that pandas defines but that would silently ignore weights are overridden rather than inherited, so a method either returns a weighted result or warns that it cannot. +The package's central design choice is that the weight is a property of the data structure rather than an argument to a function. `MicroSeries` and `MicroDataFrame` subclass the pandas [@mckinney2010pandas; @pandas2020] structures and carry a weight vector through the operations an analysis pipeline performs. Selection, merging, grouping, reindexing, dropping and type conversion are overridden so that the weight follows the rows it describes, and aggregations that pandas defines but that would silently ignore weights are overridden rather than inherited, so a method either returns a weighted result or warns that it cannot. On top of that foundation the package implements the estimators distributional analysis reports: quantiles by the inverse cumulative distribution function, frequency-weighted variance, the Gini coefficient from the Lorenz curve, top and bottom shares with proportional handling of records tied at the cutoff, and the Foster-Greer-Thorbecke family of poverty measures [@foster1984fgt]. @@ -40,7 +40,7 @@ The first is that the estimators themselves require care. A weighted median is n The second problem is that weights must stay aligned with the data through every transformation before the estimator runs. Building an analysis dataset means merging administrative variables onto survey records, filtering to a subpopulation, grouping by geography, reindexing after a sort. Each of these is an opportunity for the weight vector to fall out of alignment with the rows it describes, and nothing raises when it does: the pipeline completes and returns a number that is wrong by an amount nobody can see. In our experience maintaining microsimulation datasets, this is a more frequent source of error than the estimator formulas, and a harder one to detect, because the result remains plausible. -`microdf` addresses both. It makes the estimator decisions once, documents them, and tests them: the quantile estimator follows the inverse CDF definition, matching the default behaviour of R's `survey::svyquantile` [@lumley2004survey; @lumley2010complex] so that results can be checked against an established implementation; the variance treats weights as frequencies, so that with integer weights it agrees with `numpy` on the replicated sample; the top-share estimator splits the record at the cutoff proportionally, so a constant distribution returns the share it should. And it keeps the weight attached to the data across the transformations between loading a file and computing a statistic, which is what makes those estimator guarantees worth anything in a real pipeline. +`microdf` addresses both. It makes the estimator decisions once, documents them, and tests them: the quantile estimator follows the inverse CDF definition, matching the default behaviour of R's `survey::svyquantile` [@lumley2004survey; @lumley2010complex] so that results can be checked against an established implementation; the variance treats weights as frequencies, so that with integer weights it agrees with `numpy` on the replicated sample; the top-share estimator splits the record at the cutoff proportionally, so a constant distribution returns the share it should. And it carries the weight with the data across the transformations between loading a file and computing a statistic, which is what makes those estimator guarantees worth anything in a real pipeline. # State of the Field @@ -62,7 +62,9 @@ R's `survey` package is the reference implementation for design-based survey inf `MicroSeries` extends `pandas.Series` with a weight vector of equal length; `MicroDataFrame` extends `pandas.DataFrame`, holds a weight column, and exposes each column as a `MicroSeries`. -Pandas methods are classified into three groups, and the classification is what makes the guarantee hold. *Scalar* methods return a single weighted statistic and are overridden to use the weights. *Vector* methods return a series aligned to the input and carry the weights through to the result. *Agnostic* methods do not depend on weighting and are inherited unchanged. Methods that would need weighting but do not yet implement it, currently `cov` and `corr`, fall through to pandas and emit a warning rather than returning an unweighted number silently. Shape-changing operations — `merge`, `groupby`, `reset_index`, `drop`, `astype` — are overridden so the weight vector follows the rows it describes. +Pandas methods are classified into three groups. *Scalar* methods return a single weighted statistic and are overridden to use the weights. *Vector* methods return a series aligned to the input and carry the weights through to the result. *Agnostic* methods do not depend on weighting and are inherited unchanged. Methods that would need weighting but do not yet implement it, currently `cov` and `corr`, fall through to pandas and emit a warning rather than returning an unweighted number silently. Shape-changing operations — selection, `merge`, `groupby`, `reset_index`, `drop`, `astype` — are overridden so the weight vector follows the rows it describes. + +The classification is explicit rather than inherited, which is a deliberate trade: a method must be considered before it is supported, and one that has not been is not silently assumed safe. Extending the set of preserved operations is the package's main axis of ongoing work. ```python import microdf as mdf From 4af194d58553cc1c49d0ab24df70ec14690aed3c Mon Sep 17 00:00:00 2001 From: Vahid Ahmadi Date: Wed, 16 Sep 2026 10:48:41 +0100 Subject: [PATCH 03/22] Add Nikhil Woodruff as an author, and acknowledge Anthony Volk and Jason DeBacker Nikhil has 85 commits to the package, including the O(N^2) fix to weight linking in copy(), which is a substantial contribution to the software. Anthony and Jason contributed packaging, linting and CI work, which the acknowledgements now record. Co-Authored-By: Claude Opus 5 (1M context) --- CITATION.cff | 4 ++++ paper.md | 5 ++++- 2 files changed, 8 insertions(+), 1 deletion(-) diff --git a/CITATION.cff b/CITATION.cff index 50e6c4b0..8dad7b26 100644 --- a/CITATION.cff +++ b/CITATION.cff @@ -15,6 +15,10 @@ authors: given-names: Max orcid: "https://orcid.org/0000-0002-1335-8277" affiliation: "PolicyEngine, Washington, DC, United States" + - family-names: Woodruff + given-names: Nikhil + orcid: "https://orcid.org/0009-0009-5004-4910" + affiliation: "PolicyEngine, Washington, DC, United States" - family-names: Juaristi given-names: María affiliation: "PolicyEngine, Washington, DC, United States" diff --git a/paper.md b/paper.md index 71d14551..e4f51327 100644 --- a/paper.md +++ b/paper.md @@ -15,6 +15,9 @@ authors: - name: Max Ghenis orcid: 0000-0002-1335-8277 affiliation: '1' + - name: Nikhil Woodruff + orcid: 0009-0009-5004-4910 + affiliation: '1' - name: María Juaristi affiliation: '1' affiliations: @@ -95,7 +98,7 @@ The package's role is that of infrastructure: it is not the visible output of an Arnold Ventures [@arnold_ventures], NEO Philanthropy [@neo_philanthropy], the Gerald Huff Fund for Humanity, and the National Science Foundation (NSF POSE Phase I, Award 2518372) [@nsf_pose] funded this work in the US. The Nuffield Foundation has funded the UK work since September 2024 [@nuffield2024grant]. These funders had no involvement in the design, development, or content of this software or paper. All authors are employed by PolicyEngine and may benefit reputationally from the software's adoption; this relationship is disclosed here as a potential conflict of interest. -We thank all contributors to `microdf`, and Thomas Lumley, whose `survey` package provides the reference against which several of the estimators here are checked. +We thank Anthony Volk and Jason DeBacker for their contributions to the package, and all other contributors to `microdf`. We also thank Thomas Lumley, whose `survey` package provides the reference against which several of the estimators here are checked. # AI Usage Disclosure From 8f61e7a5f335b2fb7790de984bf0216dd17fce12 Mon Sep 17 00:00:00 2001 From: Vahid Ahmadi Date: Wed, 16 Sep 2026 10:58:27 +0100 Subject: [PATCH 04/22] Describe replicate-weight variance in the paper The package now estimates standard errors from replicate weights (#320), so the comparison table and the scope paragraph say what it does and does not do: replicate weights yes, variance from a design specification no. Co-Authored-By: Claude Opus 5 (1M context) --- paper.md | 14 +++++++++++--- 1 file changed, 11 insertions(+), 3 deletions(-) diff --git a/paper.md b/paper.md index e4f51327..61a31760 100644 --- a/paper.md +++ b/paper.md @@ -47,11 +47,11 @@ The second problem is that weights must stay aligned with the data through every # State of the Field -Several tools compute weighted statistics, but the combination `microdf` occupies — pandas-native structures, a distributional estimator set, and no complex-survey design object — is not otherwise filled. +Several tools compute weighted statistics, but the combination `microdf` occupies — pandas-native structures, a distributional estimator set, and replicate-weight variance without a complex-survey design object — is not otherwise filled. | Tool | Weighted quantiles | Inequality and poverty measures | pandas-native | Design-based variance | |---|---|---|---|---| -| `microdf` | Yes | Gini, top and bottom shares, FGT poverty | Yes | No | +| `microdf` | Yes | Gini, top and bottom shares, FGT poverty | Yes | Replicate weights | | `samplics` [@samplics] | Yes | No | Partly | Yes | | `statsmodels` [@seabold2010statsmodels] | Limited | No | Partly | Partly | | R `survey` [@lumley2004survey] | Yes | Limited | No (R) | Yes | @@ -59,7 +59,7 @@ Several tools compute weighted statistics, but the combination `microdf` occupie R's `survey` package is the reference implementation for design-based survey inference and remains the right tool when standard errors under a complex design are required. `samplics` brings much of that machinery to Python, again centred on sampling design. Neither is built around the inequality and poverty estimators that distributional policy analysis reports, and neither returns objects that behave like a `DataFrame` in an existing pandas pipeline. -`microdf` deliberately does not implement design-based variance estimation. Its weights are population weights: it estimates the statistic, not the sampling error around it. Analysts needing standard errors under a stratified or clustered design should use `survey` or `samplics`. The trade is a much smaller interface, and statistics that compose with the pandas code analysts already have. +`microdf` estimates variance from replicate weights rather than from a survey design specification. Given the replicate weight matrix that products such as the CPS and ACS publish, it recomputes a statistic once per replicate and scales the spread by the factor for the replication scheme, which requires no analytic formula and therefore works for the Gini coefficient and quantiles as readily as for a mean. What it does not do is derive variance from stratum and cluster identifiers, so analysts who need standard errors under a design specification, or who hold no replicate weights, should use `survey` or `samplics`. The replicate estimators also assume weights as published: weights calibrated to external targets no longer correspond to the original replication scheme. # Software Design @@ -86,6 +86,14 @@ regional = df.merge(geography, on="household_id") # weights follow the join regional.groupby("region").income.median() # and the grouping ``` +Standard errors come from replicate weights rather than an analytic formula, so they are available for every estimator rather than the few with tractable variance: + +```python +series.replicate_standard_error( + lambda s: s.gini(), replicate_weights, method="successive-difference" +) +``` + Statistics requiring a decision take it as an explicit argument rather than choosing silently: `gini` accepts a `negatives` policy, `var` takes `ddof`, and behaviour on zero-weight records is documented and tested. The estimators are small and independently checkable: weighted quantiles sort by value, accumulate weight, and return the smallest value whose cumulative weight share reaches `q`, after dropping zero-weight records so they cannot be selected; the Gini is computed from the Lorenz curve over weighted cumulative population and income; poverty measures follow the FGT family, with rate, gap, deep gap, and squared gap. # Research Impact Statement From f039188b50cceb16991941c3c841d97ead6f748f Mon Sep 17 00:00:00 2001 From: Vahid Ahmadi Date: Wed, 16 Sep 2026 11:06:19 +0100 Subject: [PATCH 05/22] Correct the paper against a compliance review Three substantive fixes. The contributor count said sixteen where the contributors API returns nine humans, and an editor can check that in one call. The description of agnostic methods contradicted the code, where the sole member of AGNOSTIC_FUNCTIONS is quantile, which is fully weighted. And the replication scale factors were four constants with no provenance, so they now cite Wolter and, for the successive-difference scheme, Fay and Train. Also: samplics is a journal article rather than software, the first example now runs as printed, the statsmodels row says what it actually offers, and two unfalsifiable lines are gone. Co-Authored-By: Claude Opus 5 (1M context) --- paper.bib | 35 ++++++++++++++++++++++++++++------- paper.md | 27 ++++++++++++++++----------- 2 files changed, 44 insertions(+), 18 deletions(-) diff --git a/paper.bib b/paper.bib index c6e092f1..d18b9ccd 100644 --- a/paper.bib +++ b/paper.bib @@ -33,7 +33,7 @@ @software{claude2026 title={{Claude}}, author={{Anthropic}}, year={2026}, - note={Opus 4 model used for code refactoring assistance}, + note={Opus model used for code refactoring, test authoring and drafting assistance}, url={https://www.anthropic.com/claude} } @@ -100,10 +100,31 @@ @book{lumley2010complex doi = {10.1002/9780470580066} } -@software{samplics, - author = {Diallo, Mamadou S.}, - title = {samplics: Sampling techniques for complex survey designs}, - year = {2021}, - url = {https://github.com/samplics-org/samplics}, - doi = {10.21105/joss.03376} +@article{samplics, + author = {Diallo, Mamadou S.}, + title = {samplics: a {P}ython Package for selecting, weighting and analyzing data from complex sampling designs}, + journal = {Journal of Open Source Software}, + volume = {6}, + number = {66}, + pages = {3376}, + year = {2021}, + doi = {10.21105/joss.03376} +} + +@book{wolter2007variance, + author = {Wolter, Kirk M.}, + title = {Introduction to Variance Estimation}, + edition = {2}, + publisher = {Springer}, + year = {2007}, + doi = {10.1007/978-0-387-35099-8} +} + +@inproceedings{fay1995successive, + author = {Fay, Robert E. and Train, George F.}, + title = {Aspects of Survey and Model-Based Postcensal Estimation of Income and Poverty Characteristics for States and Counties}, + booktitle = {Proceedings of the Section on Government Statistics}, + publisher = {American Statistical Association}, + pages = {154--159}, + year = {1995} } diff --git a/paper.md b/paper.md index 61a31760..cf57ada2 100644 --- a/paper.md +++ b/paper.md @@ -43,17 +43,17 @@ The first is that the estimators themselves require care. A weighted median is n The second problem is that weights must stay aligned with the data through every transformation before the estimator runs. Building an analysis dataset means merging administrative variables onto survey records, filtering to a subpopulation, grouping by geography, reindexing after a sort. Each of these is an opportunity for the weight vector to fall out of alignment with the rows it describes, and nothing raises when it does: the pipeline completes and returns a number that is wrong by an amount nobody can see. In our experience maintaining microsimulation datasets, this is a more frequent source of error than the estimator formulas, and a harder one to detect, because the result remains plausible. -`microdf` addresses both. It makes the estimator decisions once, documents them, and tests them: the quantile estimator follows the inverse CDF definition, matching the default behaviour of R's `survey::svyquantile` [@lumley2004survey; @lumley2010complex] so that results can be checked against an established implementation; the variance treats weights as frequencies, so that with integer weights it agrees with `numpy` on the replicated sample; the top-share estimator splits the record at the cutoff proportionally, so a constant distribution returns the share it should. And it carries the weight with the data across the transformations between loading a file and computing a statistic, which is what makes those estimator guarantees worth anything in a real pipeline. +`microdf` addresses both. It makes the estimator decisions once, documents them, and tests them: the quantile estimator follows the inverse CDF definition, matching the default behaviour of R's `survey::svyquantile` [@lumley2004survey; @lumley2010complex] so that results can be checked against an established implementation; the variance treats weights as frequencies, so that with integer weights it agrees with `numpy` on the replicated sample; the top-share estimator splits the record at the cutoff proportionally, so a constant distribution returns the share it should. And it carries the weight with the data across the transformations between loading a file and computing a statistic, so those estimator guarantees still hold at the point the statistic is taken. # State of the Field -Several tools compute weighted statistics, but the combination `microdf` occupies — pandas-native structures, a distributional estimator set, and replicate-weight variance without a complex-survey design object — is not otherwise filled. +Several tools compute weighted statistics. `microdf` combines pandas-native structures, a distributional estimator set, and replicate-weight variance without requiring a complex-survey design object. | Tool | Weighted quantiles | Inequality and poverty measures | pandas-native | Design-based variance | |---|---|---|---|---| | `microdf` | Yes | Gini, top and bottom shares, FGT poverty | Yes | Replicate weights | | `samplics` [@samplics] | Yes | No | Partly | Yes | -| `statsmodels` [@seabold2010statsmodels] | Limited | No | Partly | Partly | +| `statsmodels` [@seabold2010statsmodels] | `DescrStatsW` only | No | Partly | Via its survey module | | R `survey` [@lumley2004survey] | Yes | Limited | No (R) | Yes | | pandas + manual weighting | Hand-written | Hand-written | Yes | No | @@ -65,28 +65,33 @@ R's `survey` package is the reference implementation for design-based survey inf `MicroSeries` extends `pandas.Series` with a weight vector of equal length; `MicroDataFrame` extends `pandas.DataFrame`, holds a weight column, and exposes each column as a `MicroSeries`. -Pandas methods are classified into three groups. *Scalar* methods return a single weighted statistic and are overridden to use the weights. *Vector* methods return a series aligned to the input and carry the weights through to the result. *Agnostic* methods do not depend on weighting and are inherited unchanged. Methods that would need weighting but do not yet implement it, currently `cov` and `corr`, fall through to pandas and emit a warning rather than returning an unweighted number silently. Shape-changing operations — selection, `merge`, `groupby`, `reset_index`, `drop`, `astype` — are overridden so the weight vector follows the rows it describes. +Pandas methods are classified into three groups. *Scalar* methods return a single weighted statistic and are overridden to use the weights. *Vector* methods return a series aligned to the input and carry the weights through to the result. *Agnostic* methods return either a scalar or a vector depending on their arguments, and are dispatched accordingly. Methods that would need weighting but do not yet implement it, currently `cov` and `corr`, fall through to pandas and emit a warning rather than returning an unweighted number silently. Shape-changing operations — selection, `merge`, `groupby`, `reset_index`, `drop`, `astype` — are overridden so the weight vector follows the rows it describes. -The classification is explicit rather than inherited, which is a deliberate trade: a method must be considered before it is supported, and one that has not been is not silently assumed safe. Extending the set of preserved operations is the package's main axis of ongoing work. +The classification is explicit rather than inherited, which is a deliberate trade: a method must be considered before it is supported, and one that has not been is not silently assumed safe. ```python import microdf as mdf df = mdf.MicroDataFrame( - {"income": [10_000, 30_000, 120_000], "threshold": [15_000, 15_000, 15_000]}, + { + "household_id": [1, 2, 3], + "income": [10_000, 30_000, 120_000], + "threshold": [15_000, 15_000, 15_000], + }, weights=[800, 1_200, 50], ) -df.income.median() # 30000.0, weighted +df.income.median() # 30000, weighted df.income.gini() # Lorenz-curve Gini over the weighted distribution df.income.top_10_pct_share() # proportional split at the cutoff df.poverty_rate("income", "threshold") -regional = df.merge(geography, on="household_id") # weights follow the join -regional.groupby("region").income.median() # and the grouping +# Weights follow a join and a grouping. +regional = df.merge(geography, on="household_id") +regional.groupby("region").income.median() ``` -Standard errors come from replicate weights rather than an analytic formula, so they are available for every estimator rather than the few with tractable variance: +Standard errors come from replicate weights rather than an analytic formula, so they are available for every estimator rather than the few with tractable variance. The scale applied to the spread across replicates depends on how they were constructed [@wolter2007variance], including the successive-difference scheme used for the ACS and CPS [@fay1995successive]: ```python series.replicate_standard_error( @@ -98,7 +103,7 @@ Statistics requiring a decision take it as an explicit argument rather than choo # Research Impact Statement -`microdf` has been public since June 2018, with 749 commits across 16 contributors and 11 releases. It is a dependency of both `policyengine-us` and `policyengine-uk`, and therefore sits in the computational path of PolicyEngine's published distributional estimates — the poverty rates, decile impacts, and Gini changes reported in its analyses and through its web application [@policyengine_py]. It is also used directly in standalone policy studies, including analyses of free school meals, extended childcare entitlements, and national insurance reforms. +`microdf` has been public since June 2018, with over 750 commits across nine contributors and eleven tagged releases. It is a dependency of both `policyengine-us` and `policyengine-uk`, and therefore sits in the computational path of PolicyEngine's published distributional estimates — the poverty rates, decile impacts, and Gini changes reported in its analyses and through its web application [@policyengine_py]. It is also used directly in standalone policy studies, including analyses of free school meals, extended childcare entitlements, and national insurance reforms. The package's role is that of infrastructure: it is not the visible output of an analysis, but the layer that determines whether a reported poverty rate is a population estimate or a sample artefact. Its adoption is best measured by the analyses that depend on it rather than by direct use. From 038238e93761e718505168ddf389af1ef8b75427 Mon Sep 17 00:00:00 2001 From: Vahid Ahmadi Date: Wed, 16 Sep 2026 11:54:26 +0100 Subject: [PATCH 06/22] Correct three claims the source does not support The statsmodels row credited a survey module that does not exist in the package; it has DescrStatsW for weighted statistics and nothing for design-based variance. The poverty measures were described as the Foster-Greer-Thorbecke family, but poverty_gap and squared_poverty_gap return aggregate gaps in currency units rather than the normalised indices. The paper now says so. The claim that a method either returns a weighted result or warns that it cannot was true only of cov and corr; other methods that were never overridden can still lose weights silently, which is issue #300. Also: eight distinct contributors rather than nine, once duplicate identities are merged, and the replication schemes are named precisely. Co-Authored-By: Claude Opus 5 (1M context) --- paper.bib | 19 ++++++++++++------- paper.md | 18 ++++++++---------- 2 files changed, 20 insertions(+), 17 deletions(-) diff --git a/paper.bib b/paper.bib index d18b9ccd..f6998e1b 100644 --- a/paper.bib +++ b/paper.bib @@ -1,9 +1,14 @@ -@software{policyengine_py, - title={{policyengine}}, - author={{PolicyEngine Contributors}}, - version={4.10.0}, - url={https://github.com/PolicyEngine/policyengine.py}, - year={2026} +@article{policyengine_py, + author = {Ahmadi, Vahid and Ghenis, Max and Woodruff, Nikhil and Makarchuk, Pavel and Volk, Anthony}, + title = {policyengine: A Microsimulation Tool for Tax-Benefit Policy Analysis}, + journal = {Journal of Open Source Software}, + publisher = {The Open Journal}, + volume = {11}, + number = {125}, + pages = {11115}, + year = {2026}, + doi = {10.21105/joss.11115}, + url = {https://doi.org/10.21105/joss.11115} } @misc{arnold_ventures, @@ -33,7 +38,7 @@ @software{claude2026 title={{Claude}}, author={{Anthropic}}, year={2026}, - note={Opus model used for code refactoring, test authoring and drafting assistance}, + note={Used for code refactoring, test authoring and drafting assistance}, url={https://www.anthropic.com/claude} } diff --git a/paper.md b/paper.md index cf57ada2..5a9cd917 100644 --- a/paper.md +++ b/paper.md @@ -31,15 +31,15 @@ bibliography: paper.bib `microdf` provides weighted data structures for survey microdata analysis in Python. Survey records carry sampling weights: each row stands for many households, and the weights vary by orders of magnitude within a single file. Statistics computed without them describe the sample rather than the population, and the two can differ substantially. -The package's central design choice is that the weight is a property of the data structure rather than an argument to a function. `MicroSeries` and `MicroDataFrame` subclass the pandas [@mckinney2010pandas; @pandas2020] structures and carry a weight vector through the operations an analysis pipeline performs. Selection, merging, grouping, reindexing, dropping and type conversion are overridden so that the weight follows the rows it describes, and aggregations that pandas defines but that would silently ignore weights are overridden rather than inherited, so a method either returns a weighted result or warns that it cannot. +The package's central design choice is that the weight is a property of the data structure rather than an argument to a function. `MicroSeries` and `MicroDataFrame` subclass the pandas [@mckinney2010pandas; @pandas2020] structures and carry a weight vector through the operations an analysis pipeline performs. Selection, merging, grouping, reindexing, dropping and type conversion are overridden so that the weight follows the rows it describes, and aggregations that pandas defines but that would silently ignore weights are overridden rather than inherited. -On top of that foundation the package implements the estimators distributional analysis reports: quantiles by the inverse cumulative distribution function, frequency-weighted variance, the Gini coefficient from the Lorenz curve, top and bottom shares with proportional handling of records tied at the cutoff, and the Foster-Greer-Thorbecke family of poverty measures [@foster1984fgt]. +On top of that foundation the package implements the estimators distributional analysis reports: quantiles by the inverse cumulative distribution function, frequency-weighted variance, the Gini coefficient from the Lorenz curve, top and bottom shares with proportional handling of records tied at the cutoff, and poverty measures in the Foster-Greer-Thorbecke form [@foster1984fgt]: the headcount rate, and aggregate poverty and squared-poverty gaps. # Statement of Need Analysts working with survey microdata in Python face two distinct problems, and the second is the one that causes silent errors. -The first is that the estimators themselves require care. A weighted median is not the median of the weighted values. A weighted variance requires deciding whether weights are frequencies or precision weights, and the two give different answers. A top-1% share requires deciding what happens to the records straddling the cutoff: assigning them wholly to one side introduces a bias that grows as weights grow coarser. These are solvable problems, but each is a decision that hand-written code makes implicitly, usually without recording it, so implementations diverge between analysts on exactly the edge cases that matter. +The first is in the estimators. A weighted median is not the median of the weighted values. A weighted variance requires deciding whether weights are frequencies or precision weights, and the two give different answers. A top-1% share requires deciding what happens to the records straddling the cutoff: assigning them wholly to one side introduces a bias that grows as weights grow coarser. These are solvable problems, but each is a decision that hand-written code makes implicitly, usually without recording it, so implementations diverge between analysts on exactly the edge cases that matter. The second problem is that weights must stay aligned with the data through every transformation before the estimator runs. Building an analysis dataset means merging administrative variables onto survey records, filtering to a subpopulation, grouping by geography, reindexing after a sort. Each of these is an opportunity for the weight vector to fall out of alignment with the rows it describes, and nothing raises when it does: the pipeline completes and returns a number that is wrong by an amount nobody can see. In our experience maintaining microsimulation datasets, this is a more frequent source of error than the estimator formulas, and a harder one to detect, because the result remains plausible. @@ -53,7 +53,7 @@ Several tools compute weighted statistics. `microdf` combines pandas-native stru |---|---|---|---|---| | `microdf` | Yes | Gini, top and bottom shares, FGT poverty | Yes | Replicate weights | | `samplics` [@samplics] | Yes | No | Partly | Yes | -| `statsmodels` [@seabold2010statsmodels] | `DescrStatsW` only | No | Partly | Via its survey module | +| `statsmodels` [@seabold2010statsmodels] | `DescrStatsW` only | No | Partly | No | | R `survey` [@lumley2004survey] | Yes | Limited | No (R) | Yes | | pandas + manual weighting | Hand-written | Hand-written | Yes | No | @@ -91,7 +91,7 @@ regional = df.merge(geography, on="household_id") regional.groupby("region").income.median() ``` -Standard errors come from replicate weights rather than an analytic formula, so they are available for every estimator rather than the few with tractable variance. The scale applied to the spread across replicates depends on how they were constructed [@wolter2007variance], including the successive-difference scheme used for the ACS and CPS [@fay1995successive]: +Standard errors come from replicate weights rather than an analytic formula, so they are available for every estimator rather than the few with tractable variance. The scale applied to the spread across replicates depends on how they were constructed [@wolter2007variance], including the Fay-type replication with a 4/R scale used for the ACS and the CPS ASEC [@fay1995successive], which publish 80 and 160 replicates respectively: ```python series.replicate_standard_error( @@ -99,13 +99,11 @@ series.replicate_standard_error( ) ``` -Statistics requiring a decision take it as an explicit argument rather than choosing silently: `gini` accepts a `negatives` policy, `var` takes `ddof`, and behaviour on zero-weight records is documented and tested. The estimators are small and independently checkable: weighted quantiles sort by value, accumulate weight, and return the smallest value whose cumulative weight share reaches `q`, after dropping zero-weight records so they cannot be selected; the Gini is computed from the Lorenz curve over weighted cumulative population and income; poverty measures follow the FGT family, with rate, gap, deep gap, and squared gap. +Statistics requiring a decision take it as an explicit argument rather than choosing silently: `gini` accepts a `negatives` policy, `var` takes `ddof`, and behaviour on zero-weight records is documented and tested. The estimators are small and independently checkable: weighted quantiles sort by value, accumulate weight, and return the smallest value whose cumulative weight share reaches `q`, after dropping zero-weight records so they cannot be selected; the Gini is computed from the Lorenz curve over weighted cumulative population and income; poverty measures follow Foster, Greer and Thorbecke in form, reported as a headcount rate and as aggregate gaps rather than as normalised indices, with deep variants at half the threshold. # Research Impact Statement -`microdf` has been public since June 2018, with over 750 commits across nine contributors and eleven tagged releases. It is a dependency of both `policyengine-us` and `policyengine-uk`, and therefore sits in the computational path of PolicyEngine's published distributional estimates — the poverty rates, decile impacts, and Gini changes reported in its analyses and through its web application [@policyengine_py]. It is also used directly in standalone policy studies, including analyses of free school meals, extended childcare entitlements, and national insurance reforms. - -The package's role is that of infrastructure: it is not the visible output of an analysis, but the layer that determines whether a reported poverty rate is a population estimate or a sample artefact. Its adoption is best measured by the analyses that depend on it rather than by direct use. +`microdf` has been public since June 2018, with over 750 commits across eight contributors and eleven tagged releases. It is a dependency of `policyengine` [@policyengine_py], and therefore sits in the computational path of the distributional estimates that package produces — the poverty rates, decile impacts, and Gini changes reported through PolicyEngine's analyses and its web application at [policyengine.org](https://policyengine.org). It is also used directly in public policy reform analysis in the United Kingdom and the United States. # Acknowledgements @@ -115,6 +113,6 @@ We thank Anthony Volk and Jason DeBacker for their contributions to the package, # AI Usage Disclosure -The authors used generative AI tools, specifically Claude Opus by Anthropic [@claude2026], to assist with code refactoring, test authoring, and drafting of this paper. Human authors reviewed, edited, and validated all AI-assisted outputs, and made all decisions regarding estimator definitions and software design. The authors remain fully responsible for the accuracy, originality, and correctness of all submitted materials. +The authors used generative AI tools, specifically Claude by Anthropic [@claude2026], to assist with code refactoring, test authoring, and drafting of this paper. Human authors reviewed, edited, and validated all AI-assisted outputs, and made all decisions regarding estimator definitions and software design. The authors remain fully responsible for the accuracy, originality, and correctness of all submitted materials. # References From 19fe28359bcee2484622b7ce08e5ed64c82aecb1 Mon Sep 17 00:00:00 2001 From: Vahid Ahmadi Date: Wed, 16 Sep 2026 11:56:28 +0100 Subject: [PATCH 07/22] =?UTF-8?q?Add=20Mar=C3=ADa=20Juaristi's=20ORCID?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5 (1M context) --- CITATION.cff | 1 + paper.md | 1 + 2 files changed, 2 insertions(+) diff --git a/CITATION.cff b/CITATION.cff index 8dad7b26..7ddc9ec9 100644 --- a/CITATION.cff +++ b/CITATION.cff @@ -21,6 +21,7 @@ authors: affiliation: "PolicyEngine, Washington, DC, United States" - family-names: Juaristi given-names: María + orcid: "https://orcid.org/0009-0007-4946-2248" affiliation: "PolicyEngine, Washington, DC, United States" keywords: - survey statistics diff --git a/paper.md b/paper.md index 5a9cd917..68999801 100644 --- a/paper.md +++ b/paper.md @@ -19,6 +19,7 @@ authors: orcid: 0009-0009-5004-4910 affiliation: '1' - name: María Juaristi + orcid: 0009-0007-4946-2248 affiliation: '1' affiliations: - name: PolicyEngine, Washington, DC, United States From 3ad2ebb72b64fe410cd1fa8d6ae6a0e09aeeb823 Mon Sep 17 00:00:00 2001 From: vahid-ahmadi Date: Wed, 16 Sep 2026 12:19:04 +0100 Subject: [PATCH 08/22] Update author metadata --- CITATION.cff | 8 ++++---- paper.md | 6 +++--- 2 files changed, 7 insertions(+), 7 deletions(-) diff --git a/CITATION.cff b/CITATION.cff index 7ddc9ec9..42d322c4 100644 --- a/CITATION.cff +++ b/CITATION.cff @@ -15,14 +15,14 @@ authors: given-names: Max orcid: "https://orcid.org/0000-0002-1335-8277" affiliation: "PolicyEngine, Washington, DC, United States" - - family-names: Woodruff - given-names: Nikhil - orcid: "https://orcid.org/0009-0009-5004-4910" - affiliation: "PolicyEngine, Washington, DC, United States" - family-names: Juaristi given-names: María orcid: "https://orcid.org/0009-0007-4946-2248" affiliation: "PolicyEngine, Washington, DC, United States" + - family-names: Woodruff + given-names: Nikhil + orcid: "https://orcid.org/0009-0009-5004-4910" + affiliation: "PolicyEngine, Washington, DC, United States" keywords: - survey statistics - microdata diff --git a/paper.md b/paper.md index 68999801..6f3685d9 100644 --- a/paper.md +++ b/paper.md @@ -15,12 +15,12 @@ authors: - name: Max Ghenis orcid: 0000-0002-1335-8277 affiliation: '1' - - name: Nikhil Woodruff - orcid: 0009-0009-5004-4910 - affiliation: '1' - name: María Juaristi orcid: 0009-0007-4946-2248 affiliation: '1' + - name: Nikhil Woodruff + orcid: 0009-0009-5004-4910 + affiliation: '1' affiliations: - name: PolicyEngine, Washington, DC, United States index: '1' From 97a5b2c33422e536b0b4c75be8dc35ec0efbde52 Mon Sep 17 00:00:00 2001 From: vahid-ahmadi Date: Fri, 18 Sep 2026 10:53:48 +0100 Subject: [PATCH 09/22] Format paper.md code blocks ruff 0.16.7 formats Python inside markdown code blocks, which is why Lint was failing on this branch. --- paper.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/paper.md b/paper.md index 6f3685d9..6fc71a50 100644 --- a/paper.md +++ b/paper.md @@ -82,8 +82,8 @@ df = mdf.MicroDataFrame( weights=[800, 1_200, 50], ) -df.income.median() # 30000, weighted -df.income.gini() # Lorenz-curve Gini over the weighted distribution +df.income.median() # 30000, weighted +df.income.gini() # Lorenz-curve Gini over the weighted distribution df.income.top_10_pct_share() # proportional split at the cutoff df.poverty_rate("income", "threshold") From 50b10ea9b7427211ea394273b16ccb94844ab25a Mon Sep 17 00:00:00 2001 From: vahid-ahmadi Date: Fri, 18 Sep 2026 11:14:40 +0100 Subject: [PATCH 10/22] Record who contributed what, and add PyPI download figures Adds a CRediT-style paragraph to the Acknowledgements, since the submitting author is not the main contributor and a reviewer is asked explicitly whether they made major contributions. Attributions follow the commit history rather than the author order. Also updates the commit count, replaces the tagged-release count with the thirty releases actually published to PyPI, and gives the download figure with the caveat that it counts CI installs. --- paper.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/paper.md b/paper.md index 6fc71a50..34a0cc55 100644 --- a/paper.md +++ b/paper.md @@ -104,13 +104,13 @@ Statistics requiring a decision take it as an explicit argument rather than choo # Research Impact Statement -`microdf` has been public since June 2018, with over 750 commits across eight contributors and eleven tagged releases. It is a dependency of `policyengine` [@policyengine_py], and therefore sits in the computational path of the distributional estimates that package produces — the poverty rates, decile impacts, and Gini changes reported through PolicyEngine's analyses and its web application at [policyengine.org](https://policyengine.org). It is also used directly in public policy reform analysis in the United Kingdom and the United States. +`microdf` has been public since June 2018, with over 800 commits across eight contributors and thirty releases published to PyPI, where it currently averages around 2,200 downloads a day. That figure counts automated installation in continuous integration alongside direct use, so it indicates the package's place in a working toolchain rather than the size of its readership. It is a dependency of `policyengine` [@policyengine_py], and therefore sits in the computational path of the distributional estimates that package produces — the poverty rates, decile impacts, and Gini changes reported through PolicyEngine's analyses and its web application at [policyengine.org](https://policyengine.org). It is also used directly in public policy reform analysis in the United Kingdom and the United States. # Acknowledgements Arnold Ventures [@arnold_ventures], NEO Philanthropy [@neo_philanthropy], the Gerald Huff Fund for Humanity, and the National Science Foundation (NSF POSE Phase I, Award 2518372) [@nsf_pose] funded this work in the US. The Nuffield Foundation has funded the UK work since September 2024 [@nuffield2024grant]. These funders had no involvement in the design, development, or content of this software or paper. All authors are employed by PolicyEngine and may benefit reputationally from the software's adoption; this relationship is disclosed here as a potential conflict of interest. -We thank Anthony Volk and Jason DeBacker for their contributions to the package, and all other contributors to `microdf`. We also thank Thomas Lumley, whose `survey` package provides the reference against which several of the estimators here are checked. +Max Ghenis created `microdf` in 2018 and wrote most of the estimators and the weight-preserving class machinery. María Juaristi contributed extensively to the estimators, the test suite, and the release infrastructure, including the weight-preservation work described above. Nikhil Woodruff contributed to the pandas integration and the packaging. Vahid Ahmadi contributed the replicate-weight variance estimation and prepared this paper. We thank Anthony Volk and Jason DeBacker for their contributions to the package, and all other contributors to `microdf`. We also thank Thomas Lumley, whose `survey` package provides the reference against which several of the estimators here are checked. # AI Usage Disclosure From d5306be591594732bdf2f9a6389e43feb8ebc196 Mon Sep 17 00:00:00 2001 From: vahid-ahmadi Date: Fri, 18 Sep 2026 12:22:34 +0100 Subject: [PATCH 11/22] Record the Zenodo DOI The v1.5.5 release is archived at 10.5281/zenodo.22829460, with concept DOI 10.5281/zenodo.22829459 resolving to the latest version. JOSS asks for the archive DOI at submission. --- CITATION.cff | 9 +++++++++ 1 file changed, 9 insertions(+) diff --git a/CITATION.cff b/CITATION.cff index 42d322c4..265f6fb1 100644 --- a/CITATION.cff +++ b/CITATION.cff @@ -6,6 +6,15 @@ url: "https://github.com/PolicyEngine/microdf" repository-code: "https://github.com/PolicyEngine/microdf" abstract: "Weighted pandas DataFrames and Series for survey microdata, with distributional statistics including weighted quantiles, Gini, income shares and Foster-Greer-Thorbecke poverty measures." license: MIT +version: 1.5.5 +doi: 10.5281/zenodo.22829459 +identifiers: + - type: doi + value: 10.5281/zenodo.22829459 + description: "Concept DOI, resolving to the latest archived version" + - type: doi + value: 10.5281/zenodo.22829460 + description: "Archived version 1.5.5" authors: - family-names: Ahmadi given-names: Vahid From 0e61f9dad2448fdf2daded0976f23d850fce68b5 Mon Sep 17 00:00:00 2001 From: vahid-ahmadi Date: Fri, 18 Sep 2026 13:56:42 +0100 Subject: [PATCH 12/22] Describe cov and corr as implemented, and drop the trade framing Per review: #291 and #330 landed, so cov and corr are frequency-weighted on both classes and the paragraph saying they fall through to pandas with a warning is out of date. No method now returns an unweighted result behind a warning; the warnings that remain guard values and to_numpy, which deliberately hand back plain data. Takes the suggested wording for the classification sentence. --- paper.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/paper.md b/paper.md index 34a0cc55..92e3462b 100644 --- a/paper.md +++ b/paper.md @@ -66,9 +66,9 @@ R's `survey` package is the reference implementation for design-based survey inf `MicroSeries` extends `pandas.Series` with a weight vector of equal length; `MicroDataFrame` extends `pandas.DataFrame`, holds a weight column, and exposes each column as a `MicroSeries`. -Pandas methods are classified into three groups. *Scalar* methods return a single weighted statistic and are overridden to use the weights. *Vector* methods return a series aligned to the input and carry the weights through to the result. *Agnostic* methods return either a scalar or a vector depending on their arguments, and are dispatched accordingly. Methods that would need weighting but do not yet implement it, currently `cov` and `corr`, fall through to pandas and emit a warning rather than returning an unweighted number silently. Shape-changing operations — selection, `merge`, `groupby`, `reset_index`, `drop`, `astype` — are overridden so the weight vector follows the rows it describes. +Pandas methods are classified into three groups. *Scalar* methods return a single weighted statistic and are overridden to use the weights. *Vector* methods return a series aligned to the input and carry the weights through to the result. *Agnostic* methods return either a scalar or a vector depending on their arguments, and are dispatched accordingly. `cov` and `corr` are frequency-weighted on both classes, each cell of the frame matrix being the estimator applied to that pair of columns. Where a method deliberately hands back plain data, as `values` and `to_numpy` do, it warns rather than letting an unweighted result look like a weighted one. Shape-changing operations — selection, `merge`, `groupby`, `reset_index`, `drop`, `astype` — are overridden so the weight vector follows the rows it describes. -The classification is explicit rather than inherited, which is a deliberate trade: a method must be considered before it is supported, and one that has not been is not silently assumed safe. +The classification among these three groups is a deliberate step: a method must be considered before it is supported, and one that has not been is not silently assumed safe. ```python import microdf as mdf From 01bc60873a519d1366469044af13c5a31f171ea3 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Mar=C3=ADa=20Juaristi?= <127882282+juaristi22@users.noreply.github.com> Date: Fri, 18 Sep 2026 14:22:54 +0100 Subject: [PATCH 13/22] Tighten the paper's prose and refresh the impact figures Remove the "rather than" contrasts, double negatives and other machine-sounding constructions, put headings in sentence case as in the JOSS template, and correct the research impact paragraph: PyPI has 36 releases, and downloads excluding mirrors averaged about 1,700 a day over the 30 days to 17 September 2026 (pypistats). No claim about the software changes. Co-Authored-By: Claude Fable 5.1 --- paper.md | 38 +++++++++++++++++++------------------- 1 file changed, 19 insertions(+), 19 deletions(-) diff --git a/paper.md b/paper.md index 92e3462b..59b35a01 100644 --- a/paper.md +++ b/paper.md @@ -30,23 +30,23 @@ bibliography: paper.bib # Summary -`microdf` provides weighted data structures for survey microdata analysis in Python. Survey records carry sampling weights: each row stands for many households, and the weights vary by orders of magnitude within a single file. Statistics computed without them describe the sample rather than the population, and the two can differ substantially. +`microdf` provides weighted data structures for survey microdata analysis in Python. Survey records carry sampling weights: each row stands for many households, and the weights vary by orders of magnitude within a single file. Statistics that ignore the weights describe the sample, and population figures can differ from them substantially. -The package's central design choice is that the weight is a property of the data structure rather than an argument to a function. `MicroSeries` and `MicroDataFrame` subclass the pandas [@mckinney2010pandas; @pandas2020] structures and carry a weight vector through the operations an analysis pipeline performs. Selection, merging, grouping, reindexing, dropping and type conversion are overridden so that the weight follows the rows it describes, and aggregations that pandas defines but that would silently ignore weights are overridden rather than inherited. +The package's central design choice is to store the weight on the data structure itself. `MicroSeries` and `MicroDataFrame` subclass the pandas [@mckinney2010pandas; @pandas2020] structures and carry a weight vector through the operations an analysis pipeline performs. Selection, merging, grouping, reindexing, dropping and type conversion are overridden so that the weight follows the rows it describes, and the aggregations pandas defines are overridden to use it. -On top of that foundation the package implements the estimators distributional analysis reports: quantiles by the inverse cumulative distribution function, frequency-weighted variance, the Gini coefficient from the Lorenz curve, top and bottom shares with proportional handling of records tied at the cutoff, and poverty measures in the Foster-Greer-Thorbecke form [@foster1984fgt]: the headcount rate, and aggregate poverty and squared-poverty gaps. +On that foundation the package implements the estimators that distributional analysis reports: quantiles by the inverse cumulative distribution function, frequency-weighted variance, the Gini coefficient from the Lorenz curve, top and bottom shares with proportional handling of records tied at the cutoff, and the Foster-Greer-Thorbecke poverty measures [@foster1984fgt], reported as a headcount rate and as aggregate poverty and squared-poverty gaps. -# Statement of Need +# Statement of need -Analysts working with survey microdata in Python face two distinct problems, and the second is the one that causes silent errors. +Analysts working with survey microdata in Python face two problems. The second causes silent errors. -The first is in the estimators. A weighted median is not the median of the weighted values. A weighted variance requires deciding whether weights are frequencies or precision weights, and the two give different answers. A top-1% share requires deciding what happens to the records straddling the cutoff: assigning them wholly to one side introduces a bias that grows as weights grow coarser. These are solvable problems, but each is a decision that hand-written code makes implicitly, usually without recording it, so implementations diverge between analysts on exactly the edge cases that matter. +The first is in the estimators. A weighted median is not the median of the weighted values. A weighted variance requires deciding whether weights are frequencies or precision weights, and the two give different answers. A top-1% share requires deciding what happens to the records straddling the cutoff: assigning them wholly to one side introduces a bias that grows as weights grow coarser. Each is a decision that hand-written code makes implicitly and rarely records, so implementations diverge on exactly the edge cases that matter. -The second problem is that weights must stay aligned with the data through every transformation before the estimator runs. Building an analysis dataset means merging administrative variables onto survey records, filtering to a subpopulation, grouping by geography, reindexing after a sort. Each of these is an opportunity for the weight vector to fall out of alignment with the rows it describes, and nothing raises when it does: the pipeline completes and returns a number that is wrong by an amount nobody can see. In our experience maintaining microsimulation datasets, this is a more frequent source of error than the estimator formulas, and a harder one to detect, because the result remains plausible. +The second problem is that weights must stay aligned with the data through every transformation before the estimator runs. Building an analysis dataset means merging administrative variables onto survey records, filtering to a subpopulation, grouping by geography, reindexing after a sort. Each of these can leave the weight vector misaligned with the rows it describes, and nothing raises when it does: the pipeline completes and returns a plausible wrong number. In our experience maintaining microsimulation datasets, this is a more frequent source of error than the estimator formulas, and a harder one to detect. -`microdf` addresses both. It makes the estimator decisions once, documents them, and tests them: the quantile estimator follows the inverse CDF definition, matching the default behaviour of R's `survey::svyquantile` [@lumley2004survey; @lumley2010complex] so that results can be checked against an established implementation; the variance treats weights as frequencies, so that with integer weights it agrees with `numpy` on the replicated sample; the top-share estimator splits the record at the cutoff proportionally, so a constant distribution returns the share it should. And it carries the weight with the data across the transformations between loading a file and computing a statistic, so those estimator guarantees still hold at the point the statistic is taken. +`microdf` addresses both. It makes the estimator decisions once, documents them and tests them: the quantile estimator follows the inverse CDF definition and matches the default behaviour of R's `survey::svyquantile` [@lumley2004survey; @lumley2010complex], so results can be checked against an established implementation; the variance treats weights as frequencies, so with integer weights it agrees with `numpy` on the replicated sample; the top-share estimator splits the record at the cutoff proportionally, so a constant distribution returns the share it should. It also carries the weight with the data through every transformation between loading a file and computing a statistic, so those guarantees hold at the point the statistic is taken. -# State of the Field +# State of the field Several tools compute weighted statistics. `microdf` combines pandas-native structures, a distributional estimator set, and replicate-weight variance without requiring a complex-survey design object. @@ -58,17 +58,17 @@ Several tools compute weighted statistics. `microdf` combines pandas-native stru | R `survey` [@lumley2004survey] | Yes | Limited | No (R) | Yes | | pandas + manual weighting | Hand-written | Hand-written | Yes | No | -R's `survey` package is the reference implementation for design-based survey inference and remains the right tool when standard errors under a complex design are required. `samplics` brings much of that machinery to Python, again centred on sampling design. Neither is built around the inequality and poverty estimators that distributional policy analysis reports, and neither returns objects that behave like a `DataFrame` in an existing pandas pipeline. +R's `survey` package is the reference implementation for design-based survey inference and remains the right tool when standard errors under a complex design are required. `samplics` brings much of that machinery to Python, also centred on sampling design. Neither implements the inequality and poverty estimators that distributional policy analysis reports, and neither returns objects that behave like a `DataFrame` in an existing pandas pipeline. -`microdf` estimates variance from replicate weights rather than from a survey design specification. Given the replicate weight matrix that products such as the CPS and ACS publish, it recomputes a statistic once per replicate and scales the spread by the factor for the replication scheme, which requires no analytic formula and therefore works for the Gini coefficient and quantiles as readily as for a mean. What it does not do is derive variance from stratum and cluster identifiers, so analysts who need standard errors under a design specification, or who hold no replicate weights, should use `survey` or `samplics`. The replicate estimators also assume weights as published: weights calibrated to external targets no longer correspond to the original replication scheme. +`microdf` estimates variance from replicate weights. Given the replicate weight matrix that products such as the CPS and ACS publish, it recomputes a statistic once per replicate and scales the spread by the factor for the replication scheme. This needs no analytic formula, so it works for the Gini coefficient and quantiles as readily as for a mean. It does not derive variance from stratum and cluster identifiers, so analysts who need standard errors under a design specification, or who hold no replicate weights, should use `survey` or `samplics`. The replicate estimators also assume weights as published: weights calibrated to external targets no longer correspond to the original replication scheme. -# Software Design +# Software design `MicroSeries` extends `pandas.Series` with a weight vector of equal length; `MicroDataFrame` extends `pandas.DataFrame`, holds a weight column, and exposes each column as a `MicroSeries`. -Pandas methods are classified into three groups. *Scalar* methods return a single weighted statistic and are overridden to use the weights. *Vector* methods return a series aligned to the input and carry the weights through to the result. *Agnostic* methods return either a scalar or a vector depending on their arguments, and are dispatched accordingly. `cov` and `corr` are frequency-weighted on both classes, each cell of the frame matrix being the estimator applied to that pair of columns. Where a method deliberately hands back plain data, as `values` and `to_numpy` do, it warns rather than letting an unweighted result look like a weighted one. Shape-changing operations — selection, `merge`, `groupby`, `reset_index`, `drop`, `astype` — are overridden so the weight vector follows the rows it describes. +Pandas methods fall into three groups. *Scalar* methods return a single weighted statistic and are overridden to use the weights. *Vector* methods return a series aligned to the input and carry the weights through to the result. *Agnostic* methods return either a scalar or a vector depending on their arguments, and are dispatched accordingly. `cov` and `corr` are frequency-weighted on both classes; each cell of the frame matrix is the estimator applied to that pair of columns. `values` and `to_numpy` hand back plain data by design and warn that the result is unweighted. Shape-changing operations are overridden so the weight vector follows the rows it describes: selection, `merge`, `groupby`, `reset_index`, `drop` and `astype`. -The classification among these three groups is a deliberate step: a method must be considered before it is supported, and one that has not been is not silently assumed safe. +The classification is a deliberate step: each method is supported only after it has been considered. ```python import microdf as mdf @@ -92,7 +92,7 @@ regional = df.merge(geography, on="household_id") regional.groupby("region").income.median() ``` -Standard errors come from replicate weights rather than an analytic formula, so they are available for every estimator rather than the few with tractable variance. The scale applied to the spread across replicates depends on how they were constructed [@wolter2007variance], including the Fay-type replication with a 4/R scale used for the ACS and the CPS ASEC [@fay1995successive], which publish 80 and 160 replicates respectively: +Standard errors come from replicate weights, so they are available for every estimator, including those with no tractable variance formula. The scale applied to the spread across replicates depends on how the replicates were constructed [@wolter2007variance], including the Fay-type replication with a 4/R scale used for the ACS and the CPS ASEC [@fay1995successive], which publish 80 and 160 replicates respectively: ```python series.replicate_standard_error( @@ -100,11 +100,11 @@ series.replicate_standard_error( ) ``` -Statistics requiring a decision take it as an explicit argument rather than choosing silently: `gini` accepts a `negatives` policy, `var` takes `ddof`, and behaviour on zero-weight records is documented and tested. The estimators are small and independently checkable: weighted quantiles sort by value, accumulate weight, and return the smallest value whose cumulative weight share reaches `q`, after dropping zero-weight records so they cannot be selected; the Gini is computed from the Lorenz curve over weighted cumulative population and income; poverty measures follow Foster, Greer and Thorbecke in form, reported as a headcount rate and as aggregate gaps rather than as normalised indices, with deep variants at half the threshold. +Statistics that require a decision take it as an explicit argument: `gini` accepts a `negatives` policy, `var` takes `ddof`, and behaviour on zero-weight records is documented and tested. The estimators are small and independently checkable: weighted quantiles sort by value, accumulate weight, and return the smallest value whose cumulative weight share reaches `q`, after dropping zero-weight records so they cannot be selected; the Gini is computed from the Lorenz curve over weighted cumulative population and income; poverty measures follow Foster, Greer and Thorbecke in form and are reported as a headcount rate and as aggregate gaps in currency units, with deep variants at half the threshold. -# Research Impact Statement +# Research impact statement -`microdf` has been public since June 2018, with over 800 commits across eight contributors and thirty releases published to PyPI, where it currently averages around 2,200 downloads a day. That figure counts automated installation in continuous integration alongside direct use, so it indicates the package's place in a working toolchain rather than the size of its readership. It is a dependency of `policyengine` [@policyengine_py], and therefore sits in the computational path of the distributional estimates that package produces — the poverty rates, decile impacts, and Gini changes reported through PolicyEngine's analyses and its web application at [policyengine.org](https://policyengine.org). It is also used directly in public policy reform analysis in the United Kingdom and the United States. +`microdf` has been public since June 2018, with over 800 commits from eight contributors and 36 releases on PyPI, where downloads excluding mirrors averaged about 1,700 a day over the 30 days to 17 September 2026. That figure includes automated installation in continuous integration alongside direct use. It is a dependency of `policyengine` [@policyengine_py], so it sits in the computational path of the distributional estimates that package produces: the poverty rates, decile impacts and Gini changes reported through PolicyEngine's analyses and its web application at [policyengine.org](https://policyengine.org). It is also used directly in public policy reform analysis in the United Kingdom and the United States. # Acknowledgements @@ -112,7 +112,7 @@ Arnold Ventures [@arnold_ventures], NEO Philanthropy [@neo_philanthropy], the Ge Max Ghenis created `microdf` in 2018 and wrote most of the estimators and the weight-preserving class machinery. María Juaristi contributed extensively to the estimators, the test suite, and the release infrastructure, including the weight-preservation work described above. Nikhil Woodruff contributed to the pandas integration and the packaging. Vahid Ahmadi contributed the replicate-weight variance estimation and prepared this paper. We thank Anthony Volk and Jason DeBacker for their contributions to the package, and all other contributors to `microdf`. We also thank Thomas Lumley, whose `survey` package provides the reference against which several of the estimators here are checked. -# AI Usage Disclosure +# AI usage disclosure The authors used generative AI tools, specifically Claude by Anthropic [@claude2026], to assist with code refactoring, test authoring, and drafting of this paper. Human authors reviewed, edited, and validated all AI-assisted outputs, and made all decisions regarding estimator definitions and software design. The authors remain fully responsible for the accuracy, originality, and correctness of all submitted materials. From 28b4accba2553609c929010db783d41bbdf4489d Mon Sep 17 00:00:00 2001 From: vahid-ahmadi Date: Fri, 18 Sep 2026 14:29:35 +0100 Subject: [PATCH 14/22] Phrase the release count so it does not go stale PyPI is already at 37 rather than 36, because every merge now mints a release. An open-ended phrasing survives the next few. --- paper.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/paper.md b/paper.md index 59b35a01..a94d922b 100644 --- a/paper.md +++ b/paper.md @@ -104,7 +104,7 @@ Statistics that require a decision take it as an explicit argument: `gini` accep # Research impact statement -`microdf` has been public since June 2018, with over 800 commits from eight contributors and 36 releases on PyPI, where downloads excluding mirrors averaged about 1,700 a day over the 30 days to 17 September 2026. That figure includes automated installation in continuous integration alongside direct use. It is a dependency of `policyengine` [@policyengine_py], so it sits in the computational path of the distributional estimates that package produces: the poverty rates, decile impacts and Gini changes reported through PolicyEngine's analyses and its web application at [policyengine.org](https://policyengine.org). It is also used directly in public policy reform analysis in the United Kingdom and the United States. +`microdf` has been public since June 2018, with over 800 commits from eight contributors and more than 35 releases on PyPI, where downloads excluding mirrors averaged about 1,700 a day over the 30 days to 17 September 2026. That figure includes automated installation in continuous integration alongside direct use. It is a dependency of `policyengine` [@policyengine_py], so it sits in the computational path of the distributional estimates that package produces: the poverty rates, decile impacts and Gini changes reported through PolicyEngine's analyses and its web application at [policyengine.org](https://policyengine.org). It is also used directly in public policy reform analysis in the United Kingdom and the United States. # Acknowledgements From c2d98027a7476767f6ee3f634ca414d6fafc38c4 Mon Sep 17 00:00:00 2001 From: vahid-ahmadi Date: Fri, 18 Sep 2026 14:34:08 +0100 Subject: [PATCH 15/22] Transpose the comparison table Criteria as rows and tools as columns. The tool names and their citations were forcing the criterion headings to wrap into three lines each, and the one long cell now sits in a column of its own rather than stretching a row. Citations move into the column headers, so samplics, statsmodels and survey stay cited; all 15 bib entries still have a citation. --- paper.md | 13 ++++++------- 1 file changed, 6 insertions(+), 7 deletions(-) diff --git a/paper.md b/paper.md index a94d922b..70e48cc9 100644 --- a/paper.md +++ b/paper.md @@ -50,13 +50,12 @@ The second problem is that weights must stay aligned with the data through every Several tools compute weighted statistics. `microdf` combines pandas-native structures, a distributional estimator set, and replicate-weight variance without requiring a complex-survey design object. -| Tool | Weighted quantiles | Inequality and poverty measures | pandas-native | Design-based variance | -|---|---|---|---|---| -| `microdf` | Yes | Gini, top and bottom shares, FGT poverty | Yes | Replicate weights | -| `samplics` [@samplics] | Yes | No | Partly | Yes | -| `statsmodels` [@seabold2010statsmodels] | `DescrStatsW` only | No | Partly | No | -| R `survey` [@lumley2004survey] | Yes | Limited | No (R) | Yes | -| pandas + manual weighting | Hand-written | Hand-written | Yes | No | +| | `microdf` | `samplics` [@samplics] | `statsmodels` [@seabold2010statsmodels] | R `survey` [@lumley2004survey] | pandas, weighted by hand | +|---|---|---|---|---|---| +| Weighted quantiles | Yes | Yes | `DescrStatsW` only | Yes | Hand-written | +| Inequality and poverty measures | Gini, top and bottom shares, FGT poverty | No | No | Limited | Hand-written | +| pandas-native | Yes | Partly | Partly | No (R) | Yes | +| Design-based variance | Replicate weights | Yes | No | Yes | No | R's `survey` package is the reference implementation for design-based survey inference and remains the right tool when standard errors under a complex design are required. `samplics` brings much of that machinery to Python, also centred on sampling design. Neither implements the inequality and poverty estimators that distributional policy analysis reports, and neither returns objects that behave like a `DataFrame` in an existing pandas pipeline. From 0c38120c679e1a5a4ef9e77fd4a0de5481e30fa8 Mon Sep 17 00:00:00 2001 From: vahid-ahmadi Date: Fri, 18 Sep 2026 14:41:06 +0100 Subject: [PATCH 16/22] Trim the contribution note, and drop two grading adverbs The contribution note is one sentence rather than four. Removes 'extensively' from it and 'substantially' from the summary: the first graded a contribution without measuring it, and the second graded a difference the sentence can state directly. Also folds the one-line 'two problems' paragraph into the one that follows it. --- paper.md | 8 +++----- 1 file changed, 3 insertions(+), 5 deletions(-) diff --git a/paper.md b/paper.md index 70e48cc9..c3e82af1 100644 --- a/paper.md +++ b/paper.md @@ -30,7 +30,7 @@ bibliography: paper.bib # Summary -`microdf` provides weighted data structures for survey microdata analysis in Python. Survey records carry sampling weights: each row stands for many households, and the weights vary by orders of magnitude within a single file. Statistics that ignore the weights describe the sample, and population figures can differ from them substantially. +`microdf` provides weighted data structures for survey microdata analysis in Python. Survey records carry sampling weights: each row stands for many households, and the weights vary by orders of magnitude within a single file. Statistics that ignore the weights describe the sample rather than the population the sample was drawn to represent. The package's central design choice is to store the weight on the data structure itself. `MicroSeries` and `MicroDataFrame` subclass the pandas [@mckinney2010pandas; @pandas2020] structures and carry a weight vector through the operations an analysis pipeline performs. Selection, merging, grouping, reindexing, dropping and type conversion are overridden so that the weight follows the rows it describes, and the aggregations pandas defines are overridden to use it. @@ -38,9 +38,7 @@ On that foundation the package implements the estimators that distributional ana # Statement of need -Analysts working with survey microdata in Python face two problems. The second causes silent errors. - -The first is in the estimators. A weighted median is not the median of the weighted values. A weighted variance requires deciding whether weights are frequencies or precision weights, and the two give different answers. A top-1% share requires deciding what happens to the records straddling the cutoff: assigning them wholly to one side introduces a bias that grows as weights grow coarser. Each is a decision that hand-written code makes implicitly and rarely records, so implementations diverge on exactly the edge cases that matter. +Analysts working with survey microdata in Python face two problems, and the second causes silent errors. The first is in the estimators. A weighted median is not the median of the weighted values. A weighted variance requires deciding whether weights are frequencies or precision weights, and the two give different answers. A top-1% share requires deciding what happens to the records straddling the cutoff: assigning them wholly to one side introduces a bias that grows as weights grow coarser. Each is a decision that hand-written code makes implicitly and rarely records, so implementations diverge on exactly the edge cases that matter. The second problem is that weights must stay aligned with the data through every transformation before the estimator runs. Building an analysis dataset means merging administrative variables onto survey records, filtering to a subpopulation, grouping by geography, reindexing after a sort. Each of these can leave the weight vector misaligned with the rows it describes, and nothing raises when it does: the pipeline completes and returns a plausible wrong number. In our experience maintaining microsimulation datasets, this is a more frequent source of error than the estimator formulas, and a harder one to detect. @@ -109,7 +107,7 @@ Statistics that require a decision take it as an explicit argument: `gini` accep Arnold Ventures [@arnold_ventures], NEO Philanthropy [@neo_philanthropy], the Gerald Huff Fund for Humanity, and the National Science Foundation (NSF POSE Phase I, Award 2518372) [@nsf_pose] funded this work in the US. The Nuffield Foundation has funded the UK work since September 2024 [@nuffield2024grant]. These funders had no involvement in the design, development, or content of this software or paper. All authors are employed by PolicyEngine and may benefit reputationally from the software's adoption; this relationship is disclosed here as a potential conflict of interest. -Max Ghenis created `microdf` in 2018 and wrote most of the estimators and the weight-preserving class machinery. María Juaristi contributed extensively to the estimators, the test suite, and the release infrastructure, including the weight-preservation work described above. Nikhil Woodruff contributed to the pandas integration and the packaging. Vahid Ahmadi contributed the replicate-weight variance estimation and prepared this paper. We thank Anthony Volk and Jason DeBacker for their contributions to the package, and all other contributors to `microdf`. We also thank Thomas Lumley, whose `survey` package provides the reference against which several of the estimators here are checked. +Max Ghenis created `microdf` in 2018 and wrote most of the estimators and the weight-preserving class machinery; María Juaristi contributed to the estimators, the test suite and the release infrastructure, including the weight preservation described above; Nikhil Woodruff to the pandas integration and the packaging; and Vahid Ahmadi contributed the replicate-weight variance estimation and prepared this paper. We thank Anthony Volk and Jason DeBacker for their contributions to the package, and all other contributors to `microdf`. We also thank Thomas Lumley, whose `survey` package provides the reference against which several of the estimators here are checked. # AI usage disclosure From 6e7792b63b3b9e154c1e73abe0b1d76101888f2b Mon Sep 17 00:00:00 2001 From: vahid-ahmadi Date: Fri, 18 Sep 2026 14:41:58 +0100 Subject: [PATCH 17/22] Lead the impact paragraph with the dependency, not the counters The paragraph opened with commits, releases and a download rate, which are the weakest evidence it has, and reached the load-bearing claim - that PolicyEngine's published distributional estimates are computed through these estimators - in its third sentence. Reversed, and the counters compressed to one closing line. Down from 104 to 78 words. --- paper.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/paper.md b/paper.md index c3e82af1..082df5fa 100644 --- a/paper.md +++ b/paper.md @@ -101,7 +101,7 @@ Statistics that require a decision take it as an explicit argument: `gini` accep # Research impact statement -`microdf` has been public since June 2018, with over 800 commits from eight contributors and more than 35 releases on PyPI, where downloads excluding mirrors averaged about 1,700 a day over the 30 days to 17 September 2026. That figure includes automated installation in continuous integration alongside direct use. It is a dependency of `policyengine` [@policyengine_py], so it sits in the computational path of the distributional estimates that package produces: the poverty rates, decile impacts and Gini changes reported through PolicyEngine's analyses and its web application at [policyengine.org](https://policyengine.org). It is also used directly in public policy reform analysis in the United Kingdom and the United States. +`microdf` is part of the foundation PolicyEngine's microsimulation stack is built on. `policyengine` [@policyengine_py] depends on it, so the poverty rates, decile impacts and Gini changes published through PolicyEngine's analyses and at [policyengine.org](https://policyengine.org) are computed through its estimators. It is also used directly in public policy reform analysis in the United Kingdom and the United States. Public since June 2018, it has over 800 commits from eight contributors and averages around 1,700 downloads a day. # Acknowledgements From 0a62b4eec89560a2ce288324fe13fd1b2087142d Mon Sep 17 00:00:00 2001 From: vahid-ahmadi Date: Fri, 18 Sep 2026 14:43:24 +0100 Subject: [PATCH 18/22] Brace-protect package names in the bibliography The reference list rendered 'Policyengine', 'Statsmodels', 'Samplics', 'Pandas-dev/pandas' and 'with python' - the style lowercases titles and then capitalises the first letter, which is wrong for names that are lowercase by convention. Braces preserve them, and the same for Python and R where they appeared mid-title. --- paper.bib | 12 ++++++------ 1 file changed, 6 insertions(+), 6 deletions(-) diff --git a/paper.bib b/paper.bib index f6998e1b..bc9d30ba 100644 --- a/paper.bib +++ b/paper.bib @@ -1,6 +1,6 @@ @article{policyengine_py, author = {Ahmadi, Vahid and Ghenis, Max and Woodruff, Nikhil and Makarchuk, Pavel and Volk, Anthony}, - title = {policyengine: A Microsimulation Tool for Tax-Benefit Policy Analysis}, + title = {{policyengine}: A Microsimulation Tool for Tax-Benefit Policy Analysis}, journal = {Journal of Open Source Software}, publisher = {The Open Journal}, volume = {11}, @@ -51,7 +51,7 @@ @misc{nuffield2024grant } @inproceedings{mckinney2010pandas, author = {McKinney, Wes}, - title = {Data Structures for Statistical Computing in Python}, + title = {Data Structures for Statistical Computing in {P}ython}, booktitle = {Proceedings of the 9th Python in Science Conference}, pages = {56--61}, year = {2010}, @@ -60,7 +60,7 @@ @inproceedings{mckinney2010pandas @software{pandas2020, author = {{The pandas development team}}, - title = {pandas-dev/pandas: Pandas}, + title = {{pandas-dev/pandas}: {pandas}}, publisher = {Zenodo}, year = {2020}, doi = {10.5281/zenodo.3509134}, @@ -69,7 +69,7 @@ @software{pandas2020 @inproceedings{seabold2010statsmodels, author = {Seabold, Skipper and Perktold, Josef}, - title = {statsmodels: Econometric and statistical modeling with Python}, + title = {{statsmodels}: Econometric and statistical modeling with {P}ython}, booktitle = {Proceedings of the 9th Python in Science Conference}, year = {2010}, doi = {10.25080/Majora-92bf1922-011} @@ -99,7 +99,7 @@ @article{lumley2004survey @book{lumley2010complex, author = {Lumley, Thomas}, - title = {Complex Surveys: A Guide to Analysis Using R}, + title = {Complex Surveys: A Guide to Analysis Using {R}}, publisher = {John Wiley and Sons}, year = {2010}, doi = {10.1002/9780470580066} @@ -107,7 +107,7 @@ @book{lumley2010complex @article{samplics, author = {Diallo, Mamadou S.}, - title = {samplics: a {P}ython Package for selecting, weighting and analyzing data from complex sampling designs}, + title = {{samplics}: a {P}ython Package for selecting, weighting and analyzing data from complex sampling designs}, journal = {Journal of Open Source Software}, volume = {6}, number = {66}, From 5b419124f1403e2253a8fa6ecd816d8e9a8c7993 Mon Sep 17 00:00:00 2001 From: vahid-ahmadi Date: Fri, 18 Sep 2026 14:52:06 +0100 Subject: [PATCH 19/22] Act on a pre-submission JOSS review Merges main, so the branch carries the weighted MicroDataFrame cov and corr from #330 and the API reference from #324. Without it a reviewer checking out this branch would see behaviour the paper does not describe. Corrects the replication paragraph: the ACS and CPS ASEC use successive-difference replication, which is the 4/R scale the paper quotes. Fay's variant of BRR is a different scheme with a 1/(R(1-k)^2) scale, and replication.py already keeps the two apart. Refreshes the state of the field. R's convey is the closest existing equivalent to microdf's estimator set and was absent; the table said R survey had limited inequality measures, which is true of survey alone and misleading once convey exists. samplics is now archived in favour of svy, and the paper said neither. Adds what DescrStatsW does cover. Adds a paragraph placing this work against the published policyengine paper, since a reviewer will otherwise ask why the dependency is not covered there. --- paper.bib | 8 ++++++++ paper.md | 18 ++++++++++-------- 2 files changed, 18 insertions(+), 8 deletions(-) diff --git a/paper.bib b/paper.bib index bc9d30ba..e07e17fb 100644 --- a/paper.bib +++ b/paper.bib @@ -133,3 +133,11 @@ @inproceedings{fay1995successive pages = {154--159}, year = {1995} } + +@manual{convey, + title = {{convey}: Income Concentration Analysis with Complex Survey Samples}, + author = {Pessoa, Djalma and Damico, Anthony and Jacob, Guilherme}, + year = {2026}, + note = {R package version 1.0.1}, + url = {https://CRAN.R-project.org/package=convey} +} diff --git a/paper.md b/paper.md index 082df5fa..5b2d3bdf 100644 --- a/paper.md +++ b/paper.md @@ -44,20 +44,22 @@ The second problem is that weights must stay aligned with the data through every `microdf` addresses both. It makes the estimator decisions once, documents them and tests them: the quantile estimator follows the inverse CDF definition and matches the default behaviour of R's `survey::svyquantile` [@lumley2004survey; @lumley2010complex], so results can be checked against an established implementation; the variance treats weights as frequencies, so with integer weights it agrees with `numpy` on the replicated sample; the top-share estimator splits the record at the cutoff proportionally, so a constant distribution returns the share it should. It also carries the weight with the data through every transformation between loading a file and computing a statistic, so those guarantees hold at the point the statistic is taken. +`microdf` began in June 2018, before any of PolicyEngine's microsimulation packages, is installed and used independently of them, and is described here as a standalone library rather than as part of the model that depends on it [@policyengine_py]. The replicate-weight variance estimation is new since that work and is documented here for the first time. + # State of the field Several tools compute weighted statistics. `microdf` combines pandas-native structures, a distributional estimator set, and replicate-weight variance without requiring a complex-survey design object. -| | `microdf` | `samplics` [@samplics] | `statsmodels` [@seabold2010statsmodels] | R `survey` [@lumley2004survey] | pandas, weighted by hand | +| | `microdf` | R `survey` + `convey` [@lumley2004survey; @convey] | `samplics` [@samplics] | `statsmodels` [@seabold2010statsmodels] | pandas, weighted by hand | |---|---|---|---|---|---| -| Weighted quantiles | Yes | Yes | `DescrStatsW` only | Yes | Hand-written | -| Inequality and poverty measures | Gini, top and bottom shares, FGT poverty | No | No | Limited | Hand-written | -| pandas-native | Yes | Partly | Partly | No (R) | Yes | -| Design-based variance | Replicate weights | Yes | No | Yes | No | +| Weighted quantiles | Yes | Yes | Yes | `DescrStatsW` only | Hand-written | +| Inequality and poverty measures | Gini, top and bottom shares, FGT poverty | Yes | No | No | Hand-written | +| pandas-native | Yes | No (R) | Partly | Partly | Yes | +| Design-based variance | Replicate weights | Yes | Yes | No | No | -R's `survey` package is the reference implementation for design-based survey inference and remains the right tool when standard errors under a complex design are required. `samplics` brings much of that machinery to Python, also centred on sampling design. Neither implements the inequality and poverty estimators that distributional policy analysis reports, and neither returns objects that behave like a `DataFrame` in an existing pandas pipeline. +R's `survey` package is the reference implementation for design-based survey inference and remains the right tool when standard errors under a complex design are required. `convey` [@convey] adds the inequality and poverty estimators on top of it, and is the closest existing equivalent to what `microdf` provides. Both require a survey design object, and both are in R. `samplics` brings design-based inference to Python, though it is now archived in favour of `svy`; neither implements distributional estimators, and neither returns objects that behave like a `DataFrame` in an existing pandas pipeline. `statsmodels`' `DescrStatsW` covers weighted moments, quantiles and covariance, but is a statistics container rather than a data structure that survives a merge or a groupby. What `microdf` offers that these do not is the combination: distributional estimators on a weighted object that stays a `DataFrame` through the transformations that precede them. -`microdf` estimates variance from replicate weights. Given the replicate weight matrix that products such as the CPS and ACS publish, it recomputes a statistic once per replicate and scales the spread by the factor for the replication scheme. This needs no analytic formula, so it works for the Gini coefficient and quantiles as readily as for a mean. It does not derive variance from stratum and cluster identifiers, so analysts who need standard errors under a design specification, or who hold no replicate weights, should use `survey` or `samplics`. The replicate estimators also assume weights as published: weights calibrated to external targets no longer correspond to the original replication scheme. +`microdf` estimates variance from replicate weights. Given the replicate weight matrix that products such as the CPS and ACS publish, it recomputes a statistic once per replicate and scales the spread by the factor for the replication scheme. This needs no analytic formula, so it works for the Gini coefficient and quantiles as readily as for a mean. It does not derive variance from stratum and cluster identifiers, so analysts who need standard errors under a design specification, or who hold no replicate weights, should use `survey`, `convey` or `svy`. The replicate estimators also assume weights as published: weights calibrated to external targets no longer correspond to the original replication scheme. # Software design @@ -89,7 +91,7 @@ regional = df.merge(geography, on="household_id") regional.groupby("region").income.median() ``` -Standard errors come from replicate weights, so they are available for every estimator, including those with no tractable variance formula. The scale applied to the spread across replicates depends on how the replicates were constructed [@wolter2007variance], including the Fay-type replication with a 4/R scale used for the ACS and the CPS ASEC [@fay1995successive], which publish 80 and 160 replicates respectively: +Standard errors come from replicate weights, so they are available for every estimator, including those with no tractable variance formula. The scale applied to the spread across replicates depends on how the replicates were constructed [@wolter2007variance], including the successive-difference replication with a 4/R scale used for the ACS and the CPS ASEC [@fay1995successive], which publish 80 and 160 replicates respectively: ```python series.replicate_standard_error( From 89585c26d53f6b52c90a64bfbcc4e0aa7b33bab9 Mon Sep 17 00:00:00 2001 From: vahid-ahmadi Date: Fri, 18 Sep 2026 14:56:27 +0100 Subject: [PATCH 20/22] Point the citation at the current archived version Zenodo has v1.5.8 and v1.5.9 now that releases are created automatically, so the version DOI in CITATION.cff was two behind and its version field said 1.5.5. The concept DOI is unchanged and still resolves to the newest. --- CITATION.cff | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/CITATION.cff b/CITATION.cff index 265f6fb1..bd963467 100644 --- a/CITATION.cff +++ b/CITATION.cff @@ -6,15 +6,15 @@ url: "https://github.com/PolicyEngine/microdf" repository-code: "https://github.com/PolicyEngine/microdf" abstract: "Weighted pandas DataFrames and Series for survey microdata, with distributional statistics including weighted quantiles, Gini, income shares and Foster-Greer-Thorbecke poverty measures." license: MIT -version: 1.5.5 +version: 1.5.9 doi: 10.5281/zenodo.22829459 identifiers: - type: doi value: 10.5281/zenodo.22829459 description: "Concept DOI, resolving to the latest archived version" - type: doi - value: 10.5281/zenodo.22829460 - description: "Archived version 1.5.5" + value: 10.5281/zenodo.22831047 + description: "Archived version 1.5.9" authors: - family-names: Ahmadi given-names: Vahid From c34ad1135157c90271df01b32aaf127667b439bc Mon Sep 17 00:00:00 2001 From: vahid-ahmadi Date: Fri, 18 Sep 2026 15:04:21 +0100 Subject: [PATCH 21/22] Simplify the comparison table, and drop the download figure Citations move out of the column headers into the paragraph below, which already discussed every tool named. The headers were four and five lines tall, and the one long cell is now 'Yes' with the estimators listed in the prose, so no cell wraps. Removes the downloads-per-day figure. Weekday traffic averages about 2,100 and weekend about 1,300, and 1,300 installs on a Sunday for a package with little external adoption is PolicyEngine's own CI rather than users. It is real traffic but not evidence of adoption, which is the only thing it was there to show. The commit, contributor and release counts stay. --- paper.md | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/paper.md b/paper.md index 5b2d3bdf..308bc96c 100644 --- a/paper.md +++ b/paper.md @@ -50,14 +50,14 @@ The second problem is that weights must stay aligned with the data through every Several tools compute weighted statistics. `microdf` combines pandas-native structures, a distributional estimator set, and replicate-weight variance without requiring a complex-survey design object. -| | `microdf` | R `survey` + `convey` [@lumley2004survey; @convey] | `samplics` [@samplics] | `statsmodels` [@seabold2010statsmodels] | pandas, weighted by hand | +| | `microdf` | R `survey` + `convey` | `samplics` | `statsmodels` | pandas by hand | |---|---|---|---|---|---| | Weighted quantiles | Yes | Yes | Yes | `DescrStatsW` only | Hand-written | -| Inequality and poverty measures | Gini, top and bottom shares, FGT poverty | Yes | No | No | Hand-written | +| Inequality and poverty measures | Yes | Yes | No | No | Hand-written | | pandas-native | Yes | No (R) | Partly | Partly | Yes | | Design-based variance | Replicate weights | Yes | Yes | No | No | -R's `survey` package is the reference implementation for design-based survey inference and remains the right tool when standard errors under a complex design are required. `convey` [@convey] adds the inequality and poverty estimators on top of it, and is the closest existing equivalent to what `microdf` provides. Both require a survey design object, and both are in R. `samplics` brings design-based inference to Python, though it is now archived in favour of `svy`; neither implements distributional estimators, and neither returns objects that behave like a `DataFrame` in an existing pandas pipeline. `statsmodels`' `DescrStatsW` covers weighted moments, quantiles and covariance, but is a statistics container rather than a data structure that survives a merge or a groupby. What `microdf` offers that these do not is the combination: distributional estimators on a weighted object that stays a `DataFrame` through the transformations that precede them. +`microdf` implements the Gini coefficient, top and bottom income shares, and Foster-Greer-Thorbecke poverty measures. R's `survey` [@lumley2004survey] is the reference implementation for design-based survey inference and remains the right tool when standard errors under a complex design are required. `convey` [@convey] adds the same family of inequality and poverty estimators on top of it, and is the closest existing equivalent to what `microdf` provides. Both require a survey design object, and both are in R. `samplics` [@samplics] brings design-based inference to Python, though it is now archived in favour of `svy`; neither implements distributional estimators, and neither returns objects that behave like a `DataFrame` in an existing pandas pipeline. `statsmodels`' `DescrStatsW` [@seabold2010statsmodels] covers weighted moments, quantiles and covariance, but is a statistics container rather than a data structure that survives a merge or a groupby. What `microdf` offers that these do not is the combination: distributional estimators on a weighted object that stays a `DataFrame` through the transformations that precede them. `microdf` estimates variance from replicate weights. Given the replicate weight matrix that products such as the CPS and ACS publish, it recomputes a statistic once per replicate and scales the spread by the factor for the replication scheme. This needs no analytic formula, so it works for the Gini coefficient and quantiles as readily as for a mean. It does not derive variance from stratum and cluster identifiers, so analysts who need standard errors under a design specification, or who hold no replicate weights, should use `survey`, `convey` or `svy`. The replicate estimators also assume weights as published: weights calibrated to external targets no longer correspond to the original replication scheme. @@ -103,7 +103,7 @@ Statistics that require a decision take it as an explicit argument: `gini` accep # Research impact statement -`microdf` is part of the foundation PolicyEngine's microsimulation stack is built on. `policyengine` [@policyengine_py] depends on it, so the poverty rates, decile impacts and Gini changes published through PolicyEngine's analyses and at [policyengine.org](https://policyengine.org) are computed through its estimators. It is also used directly in public policy reform analysis in the United Kingdom and the United States. Public since June 2018, it has over 800 commits from eight contributors and averages around 1,700 downloads a day. +`microdf` is part of the foundation PolicyEngine's microsimulation stack is built on. `policyengine` [@policyengine_py] depends on it, so the poverty rates, decile impacts and Gini changes published through PolicyEngine's analyses and at [policyengine.org](https://policyengine.org) are computed through its estimators. It is also used directly in public policy reform analysis in the United Kingdom and the United States. Public since June 2018, it has over 800 commits from eight contributors and more than 35 releases on PyPI. # Acknowledgements From be5ed493ed71243393ca33ddbf377c04c4d78ec3 Mon Sep 17 00:00:00 2001 From: vahid-ahmadi Date: Fri, 18 Sep 2026 15:12:57 +0100 Subject: [PATCH 22/22] Open up the comparison table's row spacing The row labels wrap onto two and three lines, which left the rows crowded together. arraystretch 1.5 around the table adds about 17% to its vertical span, reset to 1.0 afterwards so nothing else is affected. --- paper.md | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/paper.md b/paper.md index 308bc96c..24fa86ed 100644 --- a/paper.md +++ b/paper.md @@ -50,6 +50,8 @@ The second problem is that weights must stay aligned with the data through every Several tools compute weighted statistics. `microdf` combines pandas-native structures, a distributional estimator set, and replicate-weight variance without requiring a complex-survey design object. +\renewcommand{\arraystretch}{1.5} + | | `microdf` | R `survey` + `convey` | `samplics` | `statsmodels` | pandas by hand | |---|---|---|---|---|---| | Weighted quantiles | Yes | Yes | Yes | `DescrStatsW` only | Hand-written | @@ -57,6 +59,8 @@ Several tools compute weighted statistics. `microdf` combines pandas-native stru | pandas-native | Yes | No (R) | Partly | Partly | Yes | | Design-based variance | Replicate weights | Yes | Yes | No | No | +\renewcommand{\arraystretch}{1.0} + `microdf` implements the Gini coefficient, top and bottom income shares, and Foster-Greer-Thorbecke poverty measures. R's `survey` [@lumley2004survey] is the reference implementation for design-based survey inference and remains the right tool when standard errors under a complex design are required. `convey` [@convey] adds the same family of inequality and poverty estimators on top of it, and is the closest existing equivalent to what `microdf` provides. Both require a survey design object, and both are in R. `samplics` [@samplics] brings design-based inference to Python, though it is now archived in favour of `svy`; neither implements distributional estimators, and neither returns objects that behave like a `DataFrame` in an existing pandas pipeline. `statsmodels`' `DescrStatsW` [@seabold2010statsmodels] covers weighted moments, quantiles and covariance, but is a statistics container rather than a data structure that survives a merge or a groupby. What `microdf` offers that these do not is the combination: distributional estimators on a weighted object that stays a `DataFrame` through the transformations that precede them. `microdf` estimates variance from replicate weights. Given the replicate weight matrix that products such as the CPS and ACS publish, it recomputes a statistic once per replicate and scales the spread by the factor for the replication scheme. This needs no analytic formula, so it works for the Gini coefficient and quantiles as readily as for a mean. It does not derive variance from stratum and cluster identifiers, so analysts who need standard errors under a design specification, or who hold no replicate weights, should use `survey`, `convey` or `svy`. The replicate estimators also assume weights as published: weights calibrated to external targets no longer correspond to the original replication scheme.