Add JOSS paper for submission - #315
vahid-ahmadi wants to merge 24 commits into
Conversation
The paper is framed around the two problems microdf solves: the estimator decisions that hand-written weighting makes implicitly, and keeping weights aligned with the rows they describe through merges, filters and grouping, where a misalignment raises nothing and leaves a plausible wrong answer. Also adds a citation file, a code of conduct, and the workflow that builds a draft PDF, all of which a JOSS review checks for. The author list and ORCIDs still need confirming before submission. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The draft said a statistic is weighted whether or not the analyst remembers to weight it. Issue #300 records that weight survival is guaranteed only for the operations explicitly overridden, so the paper now says that, names them, and notes that extending the set is ongoing work. A reviewer reads the issue tracker, and a claim the tracker contradicts is worse than a narrower one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…son DeBacker Nikhil has 85 commits to the package, including the O(N^2) fix to weight linking in copy(), which is a substantial contribution to the software. Anthony and Jason contributed packaging, linting and CI work, which the acknowledgements now record. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The package now estimates standard errors from replicate weights (#320), so the comparison table and the scope paragraph say what it does and does not do: replicate weights yes, variance from a design specification no. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three substantive fixes. The contributor count said sixteen where the contributors API returns nine humans, and an editor can check that in one call. The description of agnostic methods contradicted the code, where the sole member of AGNOSTIC_FUNCTIONS is quantile, which is fully weighted. And the replication scale factors were four constants with no provenance, so they now cite Wolter and, for the successive-difference scheme, Fay and Train. Also: samplics is a journal article rather than software, the first example now runs as printed, the statsmodels row says what it actually offers, and two unfalsifiable lines are gone. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The statsmodels row credited a survey module that does not exist in the package; it has DescrStatsW for weighted statistics and nothing for design-based variance. The poverty measures were described as the Foster-Greer-Thorbecke family, but poverty_gap and squared_poverty_gap return aggregate gaps in currency units rather than the normalised indices. The paper now says so. The claim that a method either returns a weighted result or warns that it cannot was true only of cov and corr; other methods that were never overridden can still lose weights silently, which is issue #300. Also: eight distinct contributors rather than nine, once duplicate identities are merged, and the replication schemes are named precisely. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ruff 0.16.7 formats Python inside markdown code blocks, which is why Lint was failing on this branch.
Adds a CRediT-style paragraph to the Acknowledgements, since the submitting author is not the main contributor and a reviewer is asked explicitly whether they made major contributions. Attributions follow the commit history rather than the author order. Also updates the commit count, replaces the tagged-release count with the thirty releases actually published to PyPI, and gives the download figure with the caveat that it counts CI installs.
The v1.5.5 release is archived at 10.5281/zenodo.22829460, with concept DOI 10.5281/zenodo.22829459 resolving to the latest version. JOSS asks for the archive DOI at submission.
Per review: #291 and #330 landed, so cov and corr are frequency-weighted on both classes and the paragraph saying they fall through to pandas with a warning is out of date. No method now returns an unweighted result behind a warning; the warnings that remain guard values and to_numpy, which deliberately hand back plain data. Takes the suggested wording for the classification sentence.
|
Both taken, On The classification sentence is your wording verbatim. Word count 1,462, still inside 750–1,750. I have also ticked the API reference item now #324 is merged. That leaves one open item on this PR: the author list and order, which is @MaxGhenis's call rather than something I should decide. The CRediT paragraph in the Acknowledgements records who did what from the commit history, so if the order stands it is at least documented. |
Remove the "rather than" contrasts, double negatives and other machine-sounding constructions, put headings in sentence case as in the JOSS template, and correct the research impact paragraph: PyPI has 36 releases, and downloads excluding mirrors averaged about 1,700 a day over the 30 days to 17 September 2026 (pypistats). No claim about the software changes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Pushed one commit with a prose pass, since the Two figures in the research impact paragraph changed, and both are worth a look:
|
PyPI is already at 37 rather than 36, because every merge now mints a release. An open-ended phrasing survives the next few.
Criteria as rows and tools as columns. The tool names and their citations were forcing the criterion headings to wrap into three lines each, and the one long cell now sits in a column of its own rather than stretching a row. Citations move into the column headers, so samplics, statsmodels and survey stay cited; all 15 bib entries still have a citation.
The contribution note is one sentence rather than four. Removes 'extensively' from it and 'substantially' from the summary: the first graded a contribution without measuring it, and the second graded a difference the sentence can state directly. Also folds the one-line 'two problems' paragraph into the one that follows it.
The paragraph opened with commits, releases and a download rate, which are the weakest evidence it has, and reached the load-bearing claim - that PolicyEngine's published distributional estimates are computed through these estimators - in its third sentence. Reversed, and the counters compressed to one closing line. Down from 104 to 78 words.
The reference list rendered 'Policyengine', 'Statsmodels', 'Samplics', 'Pandas-dev/pandas' and 'with python' - the style lowercases titles and then capitalises the first letter, which is wrong for names that are lowercase by convention. Braces preserve them, and the same for Python and R where they appeared mid-title.
Merges main, so the branch carries the weighted MicroDataFrame cov and corr from #330 and the API reference from #324. Without it a reviewer checking out this branch would see behaviour the paper does not describe. Corrects the replication paragraph: the ACS and CPS ASEC use successive-difference replication, which is the 4/R scale the paper quotes. Fay's variant of BRR is a different scheme with a 1/(R(1-k)^2) scale, and replication.py already keeps the two apart. Refreshes the state of the field. R's convey is the closest existing equivalent to microdf's estimator set and was absent; the table said R survey had limited inequality measures, which is true of survey alone and misleading once convey exists. samplics is now archived in favour of svy, and the paper said neither. Adds what DescrStatsW does cover. Adds a paragraph placing this work against the published policyengine paper, since a reviewer will otherwise ask why the dependency is not covered there.
Zenodo has v1.5.8 and v1.5.9 now that releases are created automatically, so the version DOI in CITATION.cff was two behind and its version field said 1.5.5. The concept DOI is unchanged and still resolves to the newest.
Citations move out of the column headers into the paragraph below, which already discussed every tool named. The headers were four and five lines tall, and the one long cell is now 'Yes' with the estimators listed in the prose, so no cell wraps. Removes the downloads-per-day figure. Weekday traffic averages about 2,100 and weekend about 1,300, and 1,300 installs on a Sunday for a package with little external adoption is PolicyEngine's own CI rather than users. It is real traffic but not evidence of adoption, which is the only thing it was there to show. The commit, contributor and release counts stay.
The row labels wrap onto two and three lines, which left the rows crowded together. arraystretch 1.5 around the table adds about 17% to its vertical span, reset to 1.0 afterwards so nothing else is affected.
|
Thanks @vahid-ahmadi and @juaristi22. This got the repository into better shape than it has been in years: CI on main, tags and Zenodo, an API reference, and about a dozen correctness fixes, all of which stay. I have decided not to submit the paper. Three reasons, in order of weight. 1. The weighted layer belongs in policyengine.pypolicyengine.py is the interface, and the country packages are transitional adapters. It already treats microdf that way: 2. JOSS in 2026 is not the JOSS that took policyengineOf the 1,231 pre-review issues opened in 2026, 906 carry 3. The central claim fails on ordinary pandasThe paper says weights survive "every transformation" and that "the aggregations pandas defines are overridden". With values df.groupby("g").x.mean() # 19.0
df.groupby("g").agg({"x": "mean"}) # 55.0, no warning
df.groupby("g").agg(m=("x", "mean")) # 55.0
df.pivot_table(index="g", values="x", aggfunc=lambda z: z.mean()) # 55.0
df.apply(lambda r: r.x, axis=1).mean() # 55.0, plain Series
np.average(df.x) # 55.0
df.mean(numeric_only=True) # empty
mdf.MicroDataFrame(df).weights # all ones
pd.cut(df.x, 2).weights # all ones
(a + b).mean() != (b + a).mean() # 20.1 vs 92.9 with different weights, no warning
The estimators themselves are right: quantiles, variance, Gini, shares and cov/corr all match R What to do with this PR
Follow-ups worth filing regardless
Method: I ran an eleven-lane review with an adversarial re-check of every serious finding, plus an independent peer review from a second model, and reproduced the findings above myself against |
Summary
paper.mdandpaper.bibfor submission to the Journal of Open Source SoftwareCITATION.cff,CODE_OF_CONDUCT.mdand.github/workflows/draft-pdf.yml, all of which a JOSS review checks for and none of which this repository hadFollows the pattern of PolicyEngine/policyengine.py#264, which was accepted and published.
How the paper is framed, and why
JOSS excludes "minor utility packages" and requires substantial scholarly effort. A paper describing microdf as weighted pandas invites that objection and probably loses. This draft instead leads on the two problems the package actually solves:
survey::svyquantile, variance treats weights as frequencies so integer weights agree with numpy on the replicated sample.The State of the Field section compares against
samplics,statsmodels, Rsurveyand manual pandas.Replicate-weight variance
That comparison originally recorded a bare "No" under design-based variance, which was the weakest cell in the table. #320 adds variance and standard error estimation from replicate weights — jackknife, BRR, Fay's BRR, bootstrap and the successive-difference scheme used for the ACS and CPS — which needs no analytic formula and so works for the Gini coefficient and quantiles as readily as for a mean.
The motivation is external users rather than the table: the CPS, ACS and SIPP all publish replicate weights, so anyone adopting
microdffrom outside PolicyEngine previously had to leave the package to put a standard error on a Gini. External adoption is the weakest part of this submission, and that gap is one only external users feel.Full design-based variance stays out of scope — it needs stratum and PSU identifiers the package cannot carry, and is not meaningful for calibrated weights. The paper now says exactly that.
JOSS requirements
paper.mdwith all six required sections: Summary, Statement of Need, State of the Field, Software Design, Research Impact Statement, AI Usage Disclosure — plus Acknowledgements and Referencespaper.bib— 15 entries, all cited, no orphans, all DOIs resolveCITATION.cff, validated against schema 1.2.0CODE_OF_CONDUCT.mdWhat is left for us to do
Ordered. The first two are the ones that would sink a submission.
1. Merge #320 before submitting
The paper's comparison table, a State of the Field paragraph, and a code example all describe replicate-weight variance. That code is on the
replicate-weight-variancebranch and not onmain, so a reviewer who installs the package and runs the paper's second snippet gets anAttributeError. That fails "does the software perform the functions described in the paper?".versioning.yamlbumps and publishes on anychangelog.d/**merge, and the package is namedmicrodf-pythonon PyPI, notmicrodf. Version 1.5.2 was uploaded on 17 September, after Estimate variance from replicate weights #320 merged; installing it from PyPI and callingreplicate_standard_errorworks2. Decide authorship, and record it
A compliance review raised this independently of anything in the paper: the submitting author has ~12 of 753 commits while Max Ghenis has 583, and the JOSS reviewer checklist asks explicitly whether the submitting author made major contributions.
Confirm author list and order — @MaxGhenis, this one is yours. The submitting author is first and corresponding with 18 of 823 commits (~2%), while you created the package in June 2018 and have ~582 (~71%). The JOSS checklist asks reviewers directly whether the submitting author made major contributions, so an editor is likely to raise it. Two workable answers:
replication.pyand its tests, the paper's headline new capability — is his work, and that he prepared the paper. The CRediT sentence in the Acknowledgements records that you wrote most of the estimators and the class machinery, so the record is explicit either way.@vahid-ahmadi's preference is the second, on the grounds that you are busy and the submission workload is the larger part of the job. Whichever you choose, it needs changing in three places together:
paper.md,CITATION.cff, and the Zenodo deposit metadata in.zenodo.json— and the Zenodo record is minted per release, so the order should be settled before the release the submission cites.Either reorder, or add a CRediT-style sentence to the Acknowledgements recording who did what — a CRediT paragraph is now in the Acknowledgements, with attributions taken from the commit history rather than the author order
María Juaristi's ORCID —
0009-0007-4946-2248, verified against the ORCID registry3. Correctness issues a reviewer would run into
Merged since this PR opened: #301, #302, #303, #304 (PRs #308, #310, #309, #311), plus #305, #306 and weight-preserving serialisation (#307, #312, #313).
All three are now closed:
cov()andcorr()onMicroSeries; Weight MicroDataFrame.cov() and .corr() #330 then weighted them onMicroDataFrametoo, and the paper is updated to say sosumfunction (possible also Series) does not respectaxisarg #262 —sumdoes not respectaxis(PR Fix weighted sum axes and missing-value options #263 merged)4. Repository health, which is what a reviewer sees first
.github/workflows/master.ymlstill triggers onpush: branches: [master]while the default branch ismain. A reviewer checking "are tests run on the main branch?" sees nothingdocs/still carries_config.yml,_toc.ymlandmyst.ymlside by side, which is worth tidying but no longer breaks the builddocs/build_api.py, sharing one signature renderer with the tests, so it cannot fall behind the code or break across pandas versions5. Before hitting submit
PSLmodels/scf— the Policy Simulation Library's Survey of Consumer Finances extractor,scf/load.py,import microdf as mdf. PSL is an independent consortium, so this is the most useful of the three, though the repository was last pushed in February 2021TheAxiomFoundation/axiom-microsim— declaresmicrodfinpyproject.toml; actively developed, last pushed September 2026alimelad/policy_engile_cali_v2— an individual's PolicyEngine-derived projectCITATION.cff. Tagging had been broken since v0.4.4:.github/publish-git-tag.shcalled.github/fetch_version.py, a file that has never existed here, with|| truediscarding the error (Tag releases again #326)..zenodo.json(Add Zenodo deposit metadata #328), so the deposit credits all four authors with ORCIDs rather than whoever created the release. Verified on the record: four creators, MIT, the paper's titleA judgement still to make
Issue #314 framed this as a go/no-go rather than a drafting exercise, and that call is still @MaxGhenis's. microdf is ~1,850 lines implementing established estimators, and its adoption is overwhelmingly internal — the external dependents found so far are one consortium repository last touched in 2021 and one actively developed project. The counterweight is that it sits in the computational path of every published PolicyEngine distributional estimate. This PR exists so the text is ready if the answer is go.