microdf estimates statistics but not the sampling error around them. The State of the Field table in the JOSS paper (#315) records this as a "No" under design-based variance, against "Yes" for samplics and R survey.
Full design-based variance is out of scope: it needs stratum and PSU identifiers that microdf has no way to carry, and it is not meaningful for weights that have been calibrated to external targets, as PolicyEngine's are.
Replicate-weight variance is a different matter and is tractable. Given a matrix of replicate weights, recompute the statistic once per replicate and scale the spread by a factor determined by how the replicates were constructed:
| scheme |
factor |
| jackknife |
(R − 1) / R |
| BRR |
1 / R |
| Fay's BRR |
1 / (R(1 − k)²) |
| bootstrap |
1 / R |
| successive difference (ACS, CPS) |
4 / R |
This needs no analytic formula, so it works for every estimator in the package — including the Gini coefficient and quantiles, where the analytic variance is awkward and is precisely why users reach for R.
Why it is worth adding
The packages that would adopt microdf from outside PolicyEngine work with the CPS, ACS and SIPP, all of which publish replicate weights as standard. A user with those files currently has to leave microdf to get a standard error. That is the gap the feature closes, and external adoption is the weakest part of the JOSS case in #315.
The caveat that must be documented
It is valid only for replicate weights as published with a survey. Once weights are calibrated or reweighted to targets the original replication scheme no longer describes the estimator's variance, so this must not be applied to the enhanced FRS or CPS datasets. The docstrings say so; it is worth repeating wherever it is documented.
microdfestimates statistics but not the sampling error around them. The State of the Field table in the JOSS paper (#315) records this as a "No" under design-based variance, against "Yes" forsamplicsand Rsurvey.Full design-based variance is out of scope: it needs stratum and PSU identifiers that
microdfhas no way to carry, and it is not meaningful for weights that have been calibrated to external targets, as PolicyEngine's are.Replicate-weight variance is a different matter and is tractable. Given a matrix of replicate weights, recompute the statistic once per replicate and scale the spread by a factor determined by how the replicates were constructed:
This needs no analytic formula, so it works for every estimator in the package — including the Gini coefficient and quantiles, where the analytic variance is awkward and is precisely why users reach for R.
Why it is worth adding
The packages that would adopt
microdffrom outside PolicyEngine work with the CPS, ACS and SIPP, all of which publish replicate weights as standard. A user with those files currently has to leavemicrodfto get a standard error. That is the gap the feature closes, and external adoption is the weakest part of the JOSS case in #315.The caveat that must be documented
It is valid only for replicate weights as published with a survey. Once weights are calibrated or reweighted to targets the original replication scheme no longer describes the estimator's variance, so this must not be applied to the enhanced FRS or CPS datasets. The docstrings say so; it is worth repeating wherever it is documented.