Skip to content

Add JOSS paper for submission - #315

Open
vahid-ahmadi wants to merge 24 commits into
mainfrom
joss-paper
Open

vahid-ahmadi wants to merge 24 commits into
mainfrom
joss-paper

Conversation

@vahid-ahmadi

@vahid-ahmadi vahid-ahmadi commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Adds paper.md and paper.bib for submission to the Journal of Open Source Software
  • Adds CITATION.cff, CODE_OF_CONDUCT.md and .github/workflows/draft-pdf.yml, all of which a JOSS review checks for and none of which this repository had
  • Paper is ~1,460 words, inside JOSS's 750–1,750 range

Follows the pattern of PolicyEngine/policyengine.py#264, which was accepted and published.

How the paper is framed, and why

JOSS excludes "minor utility packages" and requires substantial scholarly effort. A paper describing microdf as weighted pandas invites that objection and probably loses. This draft instead leads on the two problems the package actually solves:

  1. Estimator decisions. A weighted median is not the median of weighted values; weighted variance requires choosing between frequency and precision weights; a top-1% share requires deciding what happens to records straddling the cutoff. microdf makes each choice once, documents it, and tests it — quantiles follow the inverse CDF so they can be checked against survey::svyquantile, variance treats weights as frequencies so integer weights agree with numpy on the replicated sample.
  2. Weight preservation through transformation. Weights must stay aligned with their rows through merges, filters, grouping and reindexing before any estimator runs. When they do not, nothing raises — the pipeline completes and returns a plausible wrong number. This is the harder problem and the one the scalar/vector/agnostic classification and the overridden shape-changing methods exist to solve.

The State of the Field section compares against samplics, statsmodels, R survey and manual pandas.

Replicate-weight variance

That comparison originally recorded a bare "No" under design-based variance, which was the weakest cell in the table. #320 adds variance and standard error estimation from replicate weights — jackknife, BRR, Fay's BRR, bootstrap and the successive-difference scheme used for the ACS and CPS — which needs no analytic formula and so works for the Gini coefficient and quantiles as readily as for a mean.

The motivation is external users rather than the table: the CPS, ACS and SIPP all publish replicate weights, so anyone adopting microdf from outside PolicyEngine previously had to leave the package to put a standard error on a Gini. External adoption is the weakest part of this submission, and that gap is one only external users feel.

Full design-based variance stays out of scope — it needs stratum and PSU identifiers the package cannot carry, and is not meaningful for calibrated weights. The paper now says exactly that.

JOSS requirements

  • OSI-approved licence (MIT)
  • Public repository, browsable source
  • Issue tracker readable without registration
  • Public development history > 6 months — since June 2018, over 800 commits, eight contributors, thirty releases on PyPI
  • paper.md with all six required sections: Summary, Statement of Need, State of the Field, Software Design, Research Impact Statement, AI Usage Disclosure — plus Acknowledgements and References
  • Word count 1,452, within 750–1,750
  • paper.bib — 15 entries, all cited, no orphans, all DOIs resolve
  • Funding acknowledgement and conflict-of-interest disclosure
  • CITATION.cff, validated against schema 1.2.0
  • CODE_OF_CONDUCT.md
  • Draft PDF workflow

What is left for us to do

Ordered. The first two are the ones that would sink a submission.

1. Merge #320 before submitting

The paper's comparison table, a State of the Field paragraph, and a code example all describe replicate-weight variance. That code is on the replicate-weight-variance branch and not on main, so a reviewer who installs the package and runs the paper's second snippet gets an AttributeError. That fails "does the software perform the functions described in the paper?".

2. Decide authorship, and record it

A compliance review raised this independently of anything in the paper: the submitting author has ~12 of 753 commits while Max Ghenis has 583, and the JOSS reviewer checklist asks explicitly whether the submitting author made major contributions.

  • Confirm author list and order — @MaxGhenis, this one is yours. The submitting author is first and corresponding with 18 of 823 commits (~2%), while you created the package in June 2018 and have ~582 (~71%). The JOSS checklist asks reviewers directly whether the submitting author made major contributions, so an editor is likely to raise it. Two workable answers:

    • You take first author. Defensible on the contribution record, but JOSS expects the first author to handle the submission, the review thread and the revisions, which is a real time commitment over several weeks.
    • Order stands as it is. @vahid-ahmadi takes first and corresponding, and with it the submission, the review correspondence and the revisions. The substantive case is that the replicate-weight variance module — replication.py and its tests, the paper's headline new capability — is his work, and that he prepared the paper. The CRediT sentence in the Acknowledgements records that you wrote most of the estimators and the class machinery, so the record is explicit either way.

    @vahid-ahmadi's preference is the second, on the grounds that you are busy and the submission workload is the larger part of the job. Whichever you choose, it needs changing in three places together: paper.md, CITATION.cff, and the Zenodo deposit metadata in .zenodo.json — and the Zenodo record is minted per release, so the order should be settled before the release the submission cites.

  • Either reorder, or add a CRediT-style sentence to the Acknowledgements recording who did what — a CRediT paragraph is now in the Acknowledgements, with attributions taken from the commit history rather than the author order

  • María Juaristi's ORCID — 0009-0007-4946-2248, verified against the ORCID registry

3. Correctness issues a reviewer would run into

Merged since this PR opened: #301, #302, #303, #304 (PRs #308, #310, #309, #311), plus #305, #306 and weight-preserving serialisation (#307, #312, #313).

All three are now closed:

4. Repository health, which is what a reviewer sees first

  • CI does not run on the default branch. .github/workflows/master.yml still triggers on push: branches: [master] while the default branch is main. A reviewer checking "are tests run on the main branch?" sees nothing
  • Fix the docs build — Restore CI on main and build docs with MyST #316 moved it to MyST, and the documentation job passes on this PR. docs/ still carries _config.yml, _toc.yml and myst.yml side by side, which is worth tidying but no longer breaks the build
  • Add an API reference — Add an API reference to the documentation #324, merged. Documents 59 methods across both classes. The page is generated by a committed script, docs/build_api.py, sharing one signature renderer with the tests, so it cannot fall behind the code or break across pandas versions
  • Add a statement of need to the README — Open the README with the problem, not the description #325, merged. It now opens with the two problems the package solves, matching the paper's framing. The README is also the PyPI long description, so it reaches PyPI at the next release

5. Before hitting submit

  • Add PyPI download figures to the Research Impact Statement — around 2,200 a day, stated with the caveat that it counts CI installs alongside direct use
  • Identify any external, non-PolicyEngine dependents — JOSS's strongest impact signal, and what this submission most lacks. A code search outside the organisation finds:
    • PSLmodels/scf — the Policy Simulation Library's Survey of Consumer Finances extractor, scf/load.py, import microdf as mdf. PSL is an independent consortium, so this is the most useful of the three, though the repository was last pushed in February 2021
    • TheAxiomFoundation/axiom-microsim — declares microdf in pyproject.toml; actively developed, last pushed September 2026
    • alimelad/policy_engile_cali_v2 — an individual's PolicyEngine-derived project
    • UBI Center repositories also use it heavily, but shared founders make them a weaker independence claim; worth citing as usage rather than as external adoption
  • Tagged release deposited to Zenodo for a DOI — done. Archived at 10.5281/zenodo.22829460 (v1.5.5), with concept DOI 10.5281/zenodo.22829459 resolving to the latest version. Recorded in CITATION.cff. Tagging had been broken since v0.4.4: .github/publish-git-tag.sh called .github/fetch_version.py, a file that has never existed here, with || true discarding the error (Tag releases again #326).
    • Enable the Zenodo–GitHub integration for this repository
    • Create a GitHub release from a tag; Zenodo then mints the DOI
    • .zenodo.json (Add Zenodo deposit metadata #328), so the deposit credits all four authors with ORCIDs rather than whoever created the release. Verified on the record: four creators, MIT, the paper's title
  • Confirm conflicts of interest on the submission form (all authors employed by PolicyEngine)

A judgement still to make

Issue #314 framed this as a go/no-go rather than a drafting exercise, and that call is still @MaxGhenis's. microdf is ~1,850 lines implementing established estimators, and its adoption is overwhelmingly internal — the external dependents found so far are one consortium repository last touched in 2021 and one actively developed project. The counterweight is that it sits in the computational path of every published PolicyEngine distributional estimate. This PR exists so the text is ready if the answer is go.

The paper is framed around the two problems microdf solves: the estimator
decisions that hand-written weighting makes implicitly, and keeping weights
aligned with the rows they describe through merges, filters and grouping,
where a misalignment raises nothing and leaves a plausible wrong answer.

Also adds a citation file, a code of conduct, and the workflow that builds a
draft PDF, all of which a JOSS review checks for.

The author list and ORCIDs still need confirming before submission.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
vahid-ahmadi and others added 2 commits September 16, 2026 10:40
The draft said a statistic is weighted whether or not the analyst remembers
to weight it. Issue #300 records that weight survival is guaranteed only for
the operations explicitly overridden, so the paper now says that, names them,
and notes that extending the set is ongoing work. A reviewer reads the issue
tracker, and a claim the tracker contradicts is worse than a narrower one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…son DeBacker

Nikhil has 85 commits to the package, including the O(N^2) fix to weight
linking in copy(), which is a substantial contribution to the software.
Anthony and Jason contributed packaging, linting and CI work, which the
acknowledgements now record.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
vahid-ahmadi and others added 8 commits September 16, 2026 10:58
The package now estimates standard errors from replicate weights (#320), so
the comparison table and the scope paragraph say what it does and does not
do: replicate weights yes, variance from a design specification no.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three substantive fixes. The contributor count said sixteen where the
contributors API returns nine humans, and an editor can check that in one
call. The description of agnostic methods contradicted the code, where the
sole member of AGNOSTIC_FUNCTIONS is quantile, which is fully weighted. And
the replication scale factors were four constants with no provenance, so
they now cite Wolter and, for the successive-difference scheme, Fay and
Train.

Also: samplics is a journal article rather than software, the first example
now runs as printed, the statsmodels row says what it actually offers, and
two unfalsifiable lines are gone.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The statsmodels row credited a survey module that does not exist in the
package; it has DescrStatsW for weighted statistics and nothing for
design-based variance.

The poverty measures were described as the Foster-Greer-Thorbecke family,
but poverty_gap and squared_poverty_gap return aggregate gaps in currency
units rather than the normalised indices. The paper now says so.

The claim that a method either returns a weighted result or warns that it
cannot was true only of cov and corr; other methods that were never
overridden can still lose weights silently, which is issue #300.

Also: eight distinct contributors rather than nine, once duplicate identities
are merged, and the replication schemes are named precisely.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ruff 0.16.7 formats Python inside markdown code blocks, which is why
Lint was failing on this branch.
Adds a CRediT-style paragraph to the Acknowledgements, since the
submitting author is not the main contributor and a reviewer is asked
explicitly whether they made major contributions. Attributions follow
the commit history rather than the author order.

Also updates the commit count, replaces the tagged-release count with
the thirty releases actually published to PyPI, and gives the download
figure with the caveat that it counts CI installs.
Comment thread paper.md Outdated
Comment thread paper.md Outdated
The v1.5.5 release is archived at 10.5281/zenodo.22829460, with concept
DOI 10.5281/zenodo.22829459 resolving to the latest version. JOSS asks
for the archive DOI at submission.
Per review: #291 and #330 landed, so cov and corr are frequency-weighted
on both classes and the paragraph saying they fall through to pandas
with a warning is out of date. No method now returns an unweighted
result behind a warning; the warnings that remain guard values and
to_numpy, which deliberately hand back plain data.

Takes the suggested wording for the classification sentence.
@vahid-ahmadi

Copy link
Copy Markdown
Contributor Author

Both taken, 0e61f9d.

On cov and corr — you are right, and it needed more than deleting the clause. With #291 and #330 both landed, no method now returns an unweighted result behind a warning, so the sentence was wrong twice over. The paragraph now says they are frequency-weighted on both classes, with each cell of the frame matrix being the estimator applied to that pair of columns, and notes that the warnings which remain guard values and to_numpy — methods that deliberately hand back plain data.

The classification sentence is your wording verbatim.

Word count 1,462, still inside 750–1,750.

I have also ticked the API reference item now #324 is merged. That leaves one open item on this PR: the author list and order, which is @MaxGhenis's call rather than something I should decide. The CRediT paragraph in the Acknowledgements records who did what from the commit history, so if the order stands it is at least documented.

Remove the "rather than" contrasts, double negatives and other
machine-sounding constructions, put headings in sentence case as in
the JOSS template, and correct the research impact paragraph: PyPI
has 36 releases, and downloads excluding mirrors averaged about 1,700
a day over the 30 days to 17 September 2026 (pypistats). No claim
about the software changes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@juaristi22

Copy link
Copy Markdown
Collaborator

Pushed one commit with a prose pass, since the cov/corr paragraph and the "deliberate trade" sentence are already right at 0e61f9d. It removes the "rather than" contrasts, the double negatives and the em-dash asides, and puts headings in sentence case as in the JOSS template. No claim about the software changes; each sentence keeps its meaning.

Two figures in the research impact paragraph changed, and both are worth a look:

  • Releases: PyPI lists 36 for microdf-python, so "thirty" is stale, and every merge now mints one, so the number will move again by submission.
  • Downloads: the paragraph said about 2,200 a day. pypistats' overall series excluding mirrors gives 51,518 downloads over the 30 days to 17 September 2026, about 1,700 a day. If the 2,200 came from the recent endpoint, which includes mirrors, that explains the gap; I used the no-mirror figure and stated the window. Restore yours if it has a different basis.

PyPI is already at 37 rather than 36, because every merge now mints a
release. An open-ended phrasing survives the next few.
Criteria as rows and tools as columns. The tool names and their
citations were forcing the criterion headings to wrap into three lines
each, and the one long cell now sits in a column of its own rather than
stretching a row.

Citations move into the column headers, so samplics, statsmodels and
survey stay cited; all 15 bib entries still have a citation.
The contribution note is one sentence rather than four. Removes
'extensively' from it and 'substantially' from the summary: the first
graded a contribution without measuring it, and the second graded a
difference the sentence can state directly.

Also folds the one-line 'two problems' paragraph into the one that
follows it.
The paragraph opened with commits, releases and a download rate, which
are the weakest evidence it has, and reached the load-bearing claim -
that PolicyEngine's published distributional estimates are computed
through these estimators - in its third sentence. Reversed, and the
counters compressed to one closing line.

Down from 104 to 78 words.
The reference list rendered 'Policyengine', 'Statsmodels', 'Samplics',
'Pandas-dev/pandas' and 'with python' - the style lowercases titles and
then capitalises the first letter, which is wrong for names that are
lowercase by convention. Braces preserve them, and the same for Python
and R where they appeared mid-title.
Merges main, so the branch carries the weighted MicroDataFrame cov and
corr from #330 and the API reference from #324. Without it a reviewer
checking out this branch would see behaviour the paper does not
describe.

Corrects the replication paragraph: the ACS and CPS ASEC use
successive-difference replication, which is the 4/R scale the paper
quotes. Fay's variant of BRR is a different scheme with a
1/(R(1-k)^2) scale, and replication.py already keeps the two apart.

Refreshes the state of the field. R's convey is the closest existing
equivalent to microdf's estimator set and was absent; the table said R
survey had limited inequality measures, which is true of survey alone
and misleading once convey exists. samplics is now archived in favour
of svy, and the paper said neither. Adds what DescrStatsW does cover.

Adds a paragraph placing this work against the published policyengine
paper, since a reviewer will otherwise ask why the dependency is not
covered there.
Zenodo has v1.5.8 and v1.5.9 now that releases are created
automatically, so the version DOI in CITATION.cff was two behind and its
version field said 1.5.5. The concept DOI is unchanged and still
resolves to the newest.
Citations move out of the column headers into the paragraph below, which
already discussed every tool named. The headers were four and five lines
tall, and the one long cell is now 'Yes' with the estimators listed in
the prose, so no cell wraps.

Removes the downloads-per-day figure. Weekday traffic averages about
2,100 and weekend about 1,300, and 1,300 installs on a Sunday for a
package with little external adoption is PolicyEngine's own CI rather
than users. It is real traffic but not evidence of adoption, which is
the only thing it was there to show. The commit, contributor and release
counts stay.
The row labels wrap onto two and three lines, which left the rows
crowded together. arraystretch 1.5 around the table adds about 17% to
its vertical span, reset to 1.0 afterwards so nothing else is affected.
@MaxGhenis

Copy link
Copy Markdown
Collaborator

Thanks @vahid-ahmadi and @juaristi22. This got the repository into better shape than it has been in years: CI on main, tags and Zenodo, an API reference, and about a dozen correctness fixes, all of which stay. I have decided not to submit the paper. Three reasons, in order of weight.

1. The weighted layer belongs in policyengine.py

policyengine.py is the interface, and the country packages are transitional adapters. It already treats microdf that way: tax_benefit_models/us/model.py calls microsim.calculate(...).values[output_order], discards the engine's MicroSeries, and rebuilds each entity as MicroDataFrame(data["person"], weights="person_weight"). So policyengine.py is the one consumer that structurally needs a weighted type, and the right end state is a module inside it: pure array estimators (gini(x, w), quantile(x, w, q) via numpy 2's weights=, top shares, poverty measures) plus a small closed wrapper that only exposes weighted operations. microdf-python then winds down on the same clock as policyengine-us, kept working for the 57 PolicyEngine repositories that still import it. A JOSS paper would describe a package we plan to fold in.

2. JOSS in 2026 is not the JOSS that took policyengine

Of the 1,231 pre-review issues opened in 2026, 906 carry rejected; 39 of the 45 flagged query-scope are rejected and the other 6 are still open. The new gate is demonstrated research impact, and our statement routes through policyengine.py, whose outputs/inequality.py computes its own Gini and top shares in numpy and pins microdf_python<1.4, so it cannot import the replicate-weight feature the paper presents as new. That sentence in the paper is currently false, and it is the sentence an editor checks.

3. The central claim fails on ordinary pandas

The paper says weights survive "every transformation" and that "the aggregations pandas defines are overridden". With values [10, 100] and weights [9, 1] (weighted mean 19):

df.groupby("g").x.mean()                  # 19.0
df.groupby("g").agg({"x": "mean"})        # 55.0, no warning
df.groupby("g").agg(m=("x", "mean"))      # 55.0
df.pivot_table(index="g", values="x", aggfunc=lambda z: z.mean())  # 55.0
df.apply(lambda r: r.x, axis=1).mean()    # 55.0, plain Series
np.average(df.x)                          # 55.0
df.mean(numeric_only=True)                # empty
mdf.MicroDataFrame(df).weights            # all ones
pd.cut(df.x, 2).weights                   # all ones
(a + b).mean() != (b + a).mean()          # 20.1 vs 92.9 with different weights, no warning

df.dropna() raises on an all-numeric frame. #264 is closed and its repro still fails. All 852 tests pass, so none of these paths is tested. A reviewer ticking "have the functional claims been confirmed" finds the first one in minutes.

The estimators themselves are right: quantiles, variance, Gini, shares and cov/corr all match R survey 4.4, laeken and numpy on the replicated sample, and the replicate-weight scale factors reproduce svrepdesign. The problem is the surface area a pandas subclass inherits, and the fix is to fail closed.

What to do with this PR

  • Split CITATION.cff and CODE_OF_CONDUCT.md into a small PR; both are useful without a paper. The Zenodo concept DOI (10.5281/zenodo.22829459) already makes the package citable.
  • Close this one. paper.md, paper.bib and draft-pdf.yml stay in the branch history if we ever revisit.

Follow-ups worth filing regardless

  • microdf: make agg (dict, named, callable), pivot_table with a callable, row-wise apply, numeric_only, the copy constructor and pd.cut either weighted or raising; add the repros above as tests; reopen groupby().agg() silently ignores weights, producing incorrect results #264.
  • microdf: docs/examples.md line 44 still says cov()/corr() are unweighted; the code and the paper say otherwise.
  • microdf: squared_poverty_gap is in currency squared and its docstring calls it "the poverty severity index"; the poverty estimators have no tests.
  • policyengine.py: replace the private _gini and the weight_fractions > 0.9 top-share cutoff with the shared estimator, so published top shares use one tie rule; lift the <1.4 pin.

Method: I ran an eleven-lane review with an adversarial re-check of every serious finding, plus an independent peer review from a second model, and reproduced the findings above myself against be5ed49 on pandas 3.0.6 before writing this.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants