Skip to content

Fix pre-submission imputation and evaluation correctness - #219

Draft
juaristi22 wants to merge 25 commits into
mainfrom
fix/publication-correctness
Draft

juaristi22 wants to merge 25 commits into
mainfrom
fix/publication-correctness

Conversation

@juaristi22

@juaristi22 juaristi22 commented Sep 16, 2026 •

Copy link
Copy Markdown
Collaborator

The pre-submission audit in #201 identified silent errors in preprocessing, conditional distributions, sampling, and method comparison. This change preserves donor-fitted transformations, computes probability-based scores, and prevents unsupported distribution forecasts from entering model selection.

Correctness changes

  • Reuse donor preprocessing on receivers and future predictions from returned fitted models; fit CV transformations on each training fold and score in original units.
  • Invert zero-inflated mixture distributions, including signed components and zero masses; apply survey weights to gates and component models.
  • Use QRF defaults that retain within-leaf variation, honor survey weights in the empirical leaf distributions, and offer independent-target marginal quantiles for distributional comparison. Sequential mode continues to support stochastic joint draws.
  • Fit requested QuantReg percentiles lazily, preserve regression intercepts for single-row/homogeneous predictions, and reject non-finite model inputs.
  • Draw OLS shocks per row with persistent independent draws and make WLS quantiles invariant to common survey-weight rescaling.
  • Score actual class probabilities, retain numeric counts as regression targets, support explicit target declarations and custom quantile grids, and handle constant or previously unseen evaluation classes.
  • Honor train_size and model seeds; select final tuned parameters through fresh training-data tuning after fold scoring.
  • Preserve Matching weights and RNG state, prune failed tuning trials, and expose failed-record counts. Matching supports donor draws; unsupported quantiles/probabilities fail explicitly and are excluded from distributional comparison.
  • Normalize mutual information using the same paired discretization and natural-log entropies; validate duplicate/overlapping columns and incompatible predictor types.
  • Run core autoimpute tests without optional R/MDN dependencies.

Fixes #202, fixes #203, fixes #204, fixes #205, fixes #206, fixes #208, fixes #209, fixes #212. Addresses the validation portion of #213.

This branch includes the existing fixes from #214 (#207), #215 (#210), and #218 (pre-existing lint failures), retaining their authorship. Those PRs can merge separately before rebasing this branch; this does not replace the JOSS paper PR.

Compatibility and publication

Numeric-coded categorical targets now require a categorical dtype or target_types={"variable": "categorical"}. Use QRF(sequential=False) for marginal quantiles of multiple targets. Quantiles of a sequential multivariate chain require a different estimator and now fail explicitly. Zero-inflated models offer independent mode for the same reason.

These corrections change imputed values and model rankings. Regenerate paper benchmarks and affected downstream datasets before submission. QRF retaining all leaf samples can increase memory use. This PR does not resolve the separate license, authorship, metadata, release/DOI, and documentation requirements tracked by #201.

Adversarial review fixes

The follow-up review reproduced five edge-case failures; all five are repaired with regression tests:

  • Promote numeric operands before subtraction in quantile loss and Matching tuning. The unsigned-count reproduction now reports loss 1.0 rather than 64.0, and tuning selects the independently verified lower-error distance.
  • Honor explicit numeric declarations for native and nullable boolean targets across fitting, comparison and cross-validation, while retaining default boolean classification.
  • Keep raw survey weights aligned and separate when the same column is transformed as a predictor, including folds, tuning, final fits and returned-model predictions.
  • Preserve constant categorical predictions with quantiles=None and attach exact point-mass probabilities using the documented result container.
  • Convert nullable numeric OLS prediction designs to native floating values while retaining intercepts and row indices.

Validation

All scheduled checks passed on repair commit b2c7ce3d8336669bc1ee1c3dfbb3a75569c6db67. The Python 3.14 CI run completed with 484 passed, 2 skipped in 239.53 seconds, including all 33 new regression cases and all 15 real-R Matching tests. The end-to-end pipeline, Python 3.12 smoke checks, lint, changelog, documentation build, and Vercel checks also passed. Formatting passed for 113 files; the post-format affected test run passed 118 tests.

The new regression suite reproduced 25 failures with 8 passing controls before repair. All 33 new cases now pass. Independent verification passed 31 checks and found all five planned repairs complete. The affected suite passed 425 tests, 1 skipped; the complete local Python 3.13.14 suite passed 469 tests, 3 skipped. The skips are the optional R and MDN runtime modules.

On a controlled normal-noise QRF fixture (3,000 training and 2,000 test observations), nominal q10/q50/q90 coverage changed from 27.45%/48.8%/70.25% to 11.15%/50.9%/88.3%. This is a regression fixture, not a universal calibration guarantee.

Optional MDN validation remains a gap because the existing workflow skips its dependency installation for this diff. This PR remains a draft; the separate submission requirements above are still open.

The CI suite reported 207 warnings, chiefly rpy2 deprecations plus solver/dependency warnings. Coverage XML was generated, but Codecov rejected the upload because a protected branch requires a token. The workflow treats that upload failure as non-blocking.

@vercel

vercel Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
microimpute-dashboard Ready Ready Preview Sep 21, 2026 11:22am UTC

@vahid-ahmadi

Copy link
Copy Markdown
Contributor

Reviewed this in depth against origin/main on identical fixtures. The science is real and I reproduced every headline correctness claim I tested — several of these were severe silent-corruption bugs and the fixes are genuine:

Claim main this PR
QRF q10/q50/q90 coverage (n=3000/2000) 18.20 / 49.80 / 80.80% 9.95 / 50.05 / 88.50%
Survey weights affect QRF median (truth 10.0) 5.448 (ignored) 9.982
OLS distinct shock values over 500 rows 10 500
OLS shock correlation between 2 targets (want ~0) 1.0 0.058
QRF imputed corr(a,b), true 0.504 0.636 (comonotonic) 0.478
Normalisation fitted on full data (leak) train only
NMI of an exact duplicate column 0.5671 1.0

Zero-inflated quantiles: 0 non-monotone entries across 13 quantiles × 300 rows, zero atom respected. QuantReg single-row constant-predictor prediction 6.9959 against a true 7. Suite: 469 passed, 3 skipped locally, matching the description; the 8 new test files contribute 161 passing tests, and 85 of them fail on main, so the suite pins real bugs.

Requesting changes on five blocking items.

Blocking

  • Cached pickled models from every prior release break at predict(). Three attribute gaps: qrf.py:379 (self.sequential), qrf.py:172 (self._weighted_leaves), ols.py:317 (self.rng). A pickle written on main and loaded on this branch gives RuntimeError: Failed to predict with QRF model: 'QRFResults' object has no attribute 'sequential', then '_QRFModel' object has no attribute '_weighted_leaves'. This is not hypothetical: policyengine-uk-data/utils/qrf.py pickles fitted models (pickle.dump at :78, pickle.load at :39) and wealth.py, vat.py, income.py and services/etb.py all reload them. Fix with getattr(..., default) or __setstate__ defaults, and add a .breaking.md saying cached models must be regenerated.
  • Matching is silently dropped from impute_all=True. autoimpute.py:307 — if model_name == best_method or model_name == "Matching": continue. 206-matching-scoring.breaking.md says Matching is excluded from distributional model selection; this also removes it from point-imputation output, which the fragment does not cover.
  • The weighted R matching path is unverified and may assign the wrong donors. statmatch_hotdeck.py:122-155 switches to RANDwNND.hotdeck with cut.don="min", which returns mtc.ids ordered by donation class, while :230-260 still assumes recipient_indices = np.arange(1, len(receiver)+1) aligns positionally. If that assumption is wrong, weighted matching silently mismatches records. Needs an R-enabled test asserting receiver→donor alignment before merge.
  • The reproducible-seed fix does not fire through the documented wiring. matching.py:24-30, :86-91, :493 — _matching_kwargs injects random_state only when self.matching_hotdeck is nnd_hotdeck_using_rpy2, but that name now resolves to the new local wrapper. A caller using the previously documented from microimpute.utils.statmatch_hotdeck import nnd_hotdeck_using_rpy2 fails the identity check and silently gets no seed — losing the reproducibility 206-models.fixed.md promises.
  • Split the PR. eabe40b is 47 files / 2,833 insertions covering eight issues in one commit, which is unbisectable if a behaviour change turns out to move the paper benchmarks the wrong way. Suggested split: (a) QRF distribution/weights/seeds, (b) zero-inflated quantiles, (c) log-loss + target types, (d) Matching, (e) preprocessing/CV. Merge Give each per-variable QRF model its own seed #214, Prune failed Matching trials instead of scoring the training mean #215 and Fix the Lint job, which is failing on main #218 separately first — authorship retention on a297d63, 4dd7caf, 2862e6f checks out, but see the note below.

Merge-order note

#214 must land before this rebases, because this PR regresses it. Its _seed_for_variable keeps only the state of #214's first commit: the sequential branch ends return self.seed + variable_offset with no modulo and no range validation (grep '2\*\*32' on the diff returns nothing), and both tuning call sites stay on self.seed. RandomForestClassifier(random_state=2**32) raises outright, so the bound is load-bearing. Only 2 of #214's 7 tests are carried.

Non-blocking

  • Weighted QRF prediction is ~10× slower — qrf.py:229-280 loops over trees in Python per row. 20k/20k, 1 target: 0.4s → 4.0s; 12 targets: 15.9s unweighted vs 49.1s weighted. Worth noting in 205.fixed.md. Conversely the PR body's "QRF retaining all leaf samples can increase memory use" is wrong in the common case — unweighted memory improved, 3091MB → 1237MB.
  • Integer count targets now return fractional values (Poisson target: integral=True → False). Downstream writes count variables such as number of children. Document in 209-target-types.breaking.md or offer an opt-out.
  • No documentation for any new public API — grep over docs/ finds nothing for target_types, create_distributional_model, QRF(sequential=, or the return_probs=True requirement, and docs/models/qrf/qrf-imputation.ipynb:63 still presents sequential as the only mode. Blocking for JOSS (Add JOSS paper for submission #201) even if not for merge.
  • Undocumented changes needing fragments: MatchingResults._predict return type changes from Dict[float, DataFrame] to DataFrame; OLSResults._predict_quantile(count_samples=) removed; validate_quantiles now rejects empty lists; log_loss now raises ValueError where it raised RuntimeError.
  • statmatch_hotdeck.py:257 fixes a real bug with no fragment — flatten() → flatten(order="F"). R fills matrices column-major, so main was interleaving recipient and donor ids across the two columns. That deserves visibility: it means prior matching output was wrong.
  • Matching is now broken in predictor analysis. predictor_analysis.py:583 passes return_probs=True unconditionally and MatchingResults._predict now raises NotImplementedError; cross_validation.py:175-180 skips Matching but _evaluate_model_performance and progressive_predictor_inclusion do not, so it dies into the blanket except at :612.
  • has_categorical is computed, used, then shadowed at autoimpute_helpers.py:318-344. The first value gates the QuantReg guard and ignores target_types; the second omits constant_targets, so a constant categorical target gets return_probs=False and then trips the new "requires predicted probabilities" error. Three call sites handle constant targets three ways.
  • Constant-target probability containers disagree: cross_validation.py:126-137 builds from the fold's singleton class, predictor_analysis.py:588-596 from the union of train and test classes, so an unseen label scores clipped-maximum log loss in one path and not the other.
  • Restate the severity in 212.fixed.md — on main an exact duplicate column scored 0.5671 rather than 1.0, and the diagonal was hard-coded to 1.0 even for constants. Anyone who used predictor analysis to choose predictors was reading corrupted numbers.
  • Confirm Implement proper error handling with logging #9 is not regressed. type_handling.py:19-21 deletes a docstring that explicitly justified treating integer {0,1} as boolean and cited that issue; a 0/1 int target is bool on main and float64 here.
  • Dead code: is_numeric_categorical_variable is unreachable from detect_variable_type but still defined and referenced by tests; cross_validation.py:492's n_jobs = 1 if model_class == Matching is dead; type_handling.py:354-359's dtype-stability check no-ops on the test path because imputer.py:593 builds a fresh DummyVariableProcessor without predictor_numeric.
  • Test quality: pytest.raises(Exception) without match= at test_autoimpute.py:379, :399, test_models/test_qrf.py:752. test_matching.py:340-341's assert all(mse < np.inf) is true by construction and never compares tuned against default. Ten new tests pass unchanged on main and pin nothing (test_qrf_distribution.py:154, test_statmatch_bridge.py:88 and :168, test_zero_inflated_quantiles.py:176, test_regression_correctness.py:35). test_qrf_distribution.py:16-20 asserts a constant against itself; test_adversarial_regressions.py:133-141 re-implements the model's own scoring formula as its oracle. Credit where due: every pytest.raises in the eight new files carries match=.

Changelog fragment types are right — the four .breaking.md files genuinely describe breaking changes. The gaps above are omissions rather than mislabelling.

@vahid-ahmadi

Copy link
Copy Markdown
Contributor

Two additions to the review above.

First, credit where it is due: the coverage figures in the description reproduce digit-for-digit. I replicated the fixture from tests/test_models/test_qrf_distribution.py:45-57 (default_rng(22), X~U(−2,2), n=3000, y=2x+N(0,1), 2000-row query, _QRFModel(42)):

this PR   q10/q50/q90:  11.15% / 50.90% / 88.30%
main      q10/q50/q90:  27.45% / 48.80% / 70.25%

Identical to the quoted numbers. On an independent fixture (normal predictors, 3000/2000) the same improvement holds — 18.20/49.80/80.80 → 9.95/50.05/88.50 — so it is not fixture-shopped. 469 passed, 3 skipped also matches the stated local figure exactly; the "484 passed, 2 skipped" CI number could not be checked here because it needs R.

Second, a concrete split, since "split the PR" is easy to say and annoying to receive without a proposal. Each step is independently revertable, in merge order:

  1. Give each per-variable QRF model its own seed #214, Prune failed Matching trials instead of scoring the training mean #215, Fix the Lint job, which is failing on main #218 as they stand — already open and reviewed; no reason to absorb them, and Give each per-variable QRF model its own seed #214 in particular must land first since this branch drops its uint32 seed bound.
  2. QRF distribution, weights and seeds — models/qrf.py, config.py defaults, test_qrf_distribution.py. Biggest downstream blast radius, and the cleanest A/B evidence.
  3. Zero-inflated quantiles — models/zero_inflated.py, test_zero_inflated_quantiles.py. Self-contained; nothing else imports the mixture inversion.
  4. Log loss and target types — comparisons/metrics.py, utils/type_handling.py, models/imputer.py validation. This is the API-breaking one and deserves its own release note.
  5. Matching — models/matching.py, utils/statmatch_hotdeck.py, the R tests. Blocked on someone with R re-running them; splitting it out unblocks everything above.
  6. Preprocessing and cross-validation — utils/data.py, evaluations/cross_validation.py, comparisons/autoimpute*.py.

If a full rebase is too costly, the minimum useful version is pulling Matching out — it is the largest single file change, and the only part with no verification path outside CI.

Two more pieces of evidence for items already listed. On integer targets, a Poisson(1.2) fixture gives unique=6 max=5 integral=True on main against unique=8 max=6 integral=False here, so count variables come back fractional. And on is_boolean_variable, the same 0/1 integer column is dtype=bool on main and float64 here — which is covered by 209-target-types.breaking.md, but the deleted docstring cited issue #9 by number as the reason for the old behaviour, so it is worth confirming #9 is not being reopened.

vahid-ahmadi added a commit that referenced this pull request Sep 21, 2026
The previous approach reformatted five documentation files and bounded
ruff to <0.17.0. Per review, the bound was a no-op against the failure
it targeted: 0.16.0 through 0.16.8 all format markdown and all satisfy
it, so only the reformatting was doing any work - and the next formatter
change inside 0.16.x would reopen the failure.

Excludes docs/**/*.md via [tool.ruff] instead, which fixes the cause
rather than the symptom, and drops the five reformatted files. That also
removes the byte-identical docs hunks that collided with #216, #217 and
#219.

The dev extra is bounded to match CI, so a contributor running
make format no longer reformats files CI then rejects.

The cause was ruff 0.16.0, not 0.16.7: 0.13, 0.14 and 0.15 report '65
files already formatted' and do not scan markdown at all.
@juaristi22

Copy link
Copy Markdown
Collaborator Author

Review of b2c7ce3: changes required. Three defects reproduced through supported public APIs, and optional MDN validation remains incomplete.

  1. Keep derived seeds within the learner's supported range. QRF target seeds and zero-inflated component seeds add offsets without bounding the result. QRF(seed=2**32-1).fit(data, ['x'], ['a', 'b'], n_estimators=2, min_samples_leaf=2) fails while fitting b with seed 4294967296; the one-target and independent-mode controls succeed. ZeroInflatedImputer(seed=2**32-1) fits an all-positive target but fails on a valid mixture of 20 zero and 40 positive observations. Bound derived seeds to uint32, retain distinct streams and independent-mode name-based derivation, and preserve Give each per-variable QRF model its own seed #214's constructor validation and tuning-stream protection during integration. Add regressions for both boundary cases.

  2. Emit exact point-mass probabilities for constant categorical targets. The new strict metric requirement breaks the documented migration path. Fit OLS or QRF with donor label='a', call predict(receiver, [0.5], return_probs=True), and pass that result to compare_metrics: both learners omit the probability container, so scoring raises a ValueError asking for the method the caller just used. Final autoimpute fitting also omits these targets from categorical detection. Emit probability 1 for the donor class and 0 for unseen evaluation classes at the common fitted-results boundary or consistently in both learners. Cover direct prediction, initial autoimpute output, and returned-model replay. Keep strict probability scoring. The get_imputations control already succeeds on truth ['a', 'b'] with loss 18.021826694558577; the supported entry points need to agree.

  3. Clear zero-inflated target-type state before refitting. ZeroInflatedImputer.fit bypasses the registry reset added to the base implementation. Reuse ZeroInflatedImputer(sequential=False), first fit numeric y, then fit the same 90 finite donor rows with target_types={'y': 'categorical'}: the second fit raises ValueError: Variable 'y' declared numeric has nonnumeric dtype. A fresh model with those second-call arguments succeeds. Reset numeric, categorical, boolean, and constant registries before type detection; add a repeat-fit test that changes declarations.

Validation gap: the current filename-based MDN gate only enables the extra when a changed path contains mdn or MDN. Shared imputer, target-typing, preprocessing, and scoring changes therefore skip both MDN test modules. Expand the dependency gate and run MDN numeric/categorical fitting, probability-based cross-validation/comparison, and transformed-input integration checks.

Integration: merge #218 first and bring #214–#217 to green CI before integrating them. Exact-head merge checks found conflicts with #214 (qrf.py, tests, changelog), #215 (matching.py, changelog), #217 (imputer and autoimpute/QRF tests), and #218 (workflow and changelog); #216 combines cleanly. Preserve the newer seed, Matching, API, and formatter repairs and rerun integrated checks. This PR remains a draft.

Validation at the reviewed head: the existing local suite passed 469 tests, with three optional runtime skips. The focused seed, mixture-seed, constant-probability, autoimpute/replay, and refit reproductions above exposed gaps beyond that suite. All eight scheduled remote checks passed, but the MDN gate means those results do not establish MDN compatibility. No combined-suite pass is claimed.

The existing review and follow-up contain additional feedback that is being investigated separately. The three defects above are the findings reproduced in this review.

vahid-ahmadi and others added 17 commits September 21, 2026 11:55
Every per-variable model builds its own generator from the seed it is
given, and all of them were handed self.seed. They therefore drew the
same random quantiles in the same row order, so variables imputed
together came out rank-comonotonic whatever their dependence in the
donor: three targets with nil conditional dependence reproduced at 0.71
Spearman, 0.11 after this change.

The offset follows the convention already used for the subsampling seed
in _apply_max_train_samples. QRF also now accepts a seed argument; there
was previously no way for a caller to vary the draws.

Fixes #207
Per review, two ways the fix could be undone silently.

_seed_for_variable fell back to offset 0 when a variable was not in
imputed_variables, which hands it the same draws as the first target -
the comonotonicity this method exists to prevent, reintroduced with no
symptom. It now raises. Today's call sites all pass post-preprocessing
names, so this is about the next refactor, not current behaviour.

Seeds were validated lazily inside _seed_for_variable, so QRF(seed=-5)
constructed fine and only failed part-way through fit, where the blanket
handler rewrapped the ValueError as RuntimeError. Validation now happens
in __init__, bools included, and the test asserts ValueError at
construction rather than RuntimeError at fit.

Also drops the docs reformatting, which belongs to #218 alone.
Both tuning handlers caught every exception and substituted the training
mean, with no log and no counter, and that score went straight into the
Optuna objective. A mean-predictor is not a neutral score - on a
low-signal target it can beat a genuine matching fit on quantile loss -
so a parameter set under which matching always failed could be selected
as best and reported as the winning method.

Both now log the exception and raise TrialPruned. The predict path keeps
its NaN fill, which is the right behaviour there, but now reports the
total number of unmatched records rather than leaving silent NaN blocks.

Fixes #210
Per review, three follow-ups.

n_failed_records kept the previous successful call's value when a
prediction raised before reaching _process_matching_results, so a caller
reading it after an exception got a stale number. It is now reset on
entry to _predict, and documented - including that Matching runs
single-threaded, so the attribute is safe in practice but would race
under concurrent calls on one fitted object.

The count is also mirrored onto result.attrs['n_failed_records'], so a
caller does not have to reach into the fitted model for it.

An all-pruned study reported only that nothing succeeded. It now carries
the most recent underlying failure, which is what a user needs when
matching fails structurally - a bad dtype or a missing R package - and
the cause was otherwise reachable only through __cause__.

Also drops the docs reformatting, which belongs to #218.
The previous commit de-indented 'raise optuna.TrialPruned() from e' out
of 'except Exception as e', so 'e' was unbound and the trial failed with
UnboundLocalError instead of pruning - which the matching failure tests
caught.
packages = ["microimpute"] with package-data "**/*" meant the
subpackages reached the wheel only as package data, swept in by a glob
that also collected whatever else was in the working tree. A wheel built
from a tree with compiled bytecode present carried 27 __pycache__
entries and around 500 KB of build-host bytecode, so wheel contents were
a function of the builder's working directory rather than the source.

setuptools.packages.find declares them properly. Verified: with 54 .pyc
files present in the tree, the built wheel now contains none, and every
subpackage imports from the installed wheel.

Also adds py.typed, so the annotations become visible to downstream
consumers; the two missing authors, which left the paper's corresponding
author out of the PyPI metadata; and classifiers and project URLs, with
no repository link previously on the PyPI page.

The absent LICENSE file is #197 and is not addressed here.

Fixes #211
Per review. Adds Python 3.12/3.13/3.14 classifiers to match
requires-python, an email for Vahid so all four authors render in
Author-email rather than one splitting into the legacy Author field, and
an explicit exclude alongside the include as belt and braces.

Rewords the changelog fragment: the bloat is latent rather than shipped.
CI builds from a fresh checkout with no imports, so published wheels
have been clean - the problem appears when anyone builds from a tree
that has been tested in. The previous wording implied released wheels
were affected.

Also drops the docs reformatting, which belongs to #218.
tests/test_autoimpute.py named Matching unconditionally whenever MDN was
absent, though HAS_MATCHING was already computed and unused. Without
rpy2 that raised NameError before the call under test ran, which is the
whole of #204: eight errors that looked like failures of the code under
test. An available_models() helper builds the list from what actually
imported.

That NameError was also masking a real problem. Four pytest.raises
calls had no match=, so they passed on any exception - including that
NameError. test_autoimpute_missing_predictors was passing on it rather
than on the missing-column error it claims to test, which is visible
now that each raises names the message it expects. A fifth test
asserted inside an except branch, so it would have passed silently if
predict ever stopped raising, and a sixth skipped its assertion when
both losses were NaN, which is exactly the case #210 produces.

ZeroInflatedImputer was reachable only by full module path despite
being a documented feature of the paper. DEFAULT_MODEL_PARAMS and
VALID_YEARS had no references anywhere in the package, and the
Imputer.fit docstring still described the bootstrap resampling scheme
that was replaced by native sample_weight support.

Fixes #204
Matching surfaces a missing predictor from R rather than from pandas, so
the message differs from the other models'. Only visible where rpy2 and
StatMatch are installed.
The DEFAULT_MODEL_PARAMS test asserted the whole mapping as a literal
against itself, which froze values nothing in the package reads - it has
no callers inside microimpute and the real defaults live in each model.
It now checks the keys and shapes, which is what a downstream caller
relies on. The constant itself stays, per @juaristi22's compatibility
fix; VALID_YEARS keeps its exact assertion because two notebooks depend
on those years.

Drops test_published_notebook_config_imports: it parsed an 8 MB notebook
to assert what the two tests above it already assert, and would fail as
a confusing KeyError if either notebook were renamed.

available_models() now returns None. autoimpute's own default is already
dependency-aware, so the helper was duplicating production logic and, as
written, only exercised Matching and MDN when MDN happened to be
installed.

The Imputer.fit weight docstring said weights go to the learner's
weighted-fit interface without noting that QuantReg and MDN raise
NotImplementedError. The paper in #201 makes claims about exactly this.

Also drops the docs reformatting, which belongs to #218.
Every per-variable model builds its own generator from the seed it is
given, and all of them were handed self.seed. They therefore drew the
same random quantiles in the same row order, so variables imputed
together came out rank-comonotonic whatever their dependence in the
donor: three targets with nil conditional dependence reproduced at 0.71
Spearman, 0.11 after this change.

The offset follows the convention already used for the subsampling seed
in _apply_max_train_samples. QRF also now accepts a seed argument; there
was previously no way for a caller to vary the draws.

Fixes #207
Both tuning handlers caught every exception and substituted the training
mean, with no log and no counter, and that score went straight into the
Optuna objective. A mean-predictor is not a neutral score - on a
low-signal target it can beat a genuine matching fit on quantile loss -
so a parameter set under which matching always failed could be selected
as best and reported as the winning method.

Both now log the exception and raise TrialPruned. The predict path keeps
its NaN fill, which is the right behaviour there, but now reports the
total number of unmatched records rather than leaving silent NaN blocks.

Fixes #210
@vahid-ahmadi

Copy link
Copy Markdown
Contributor

Re-checked the five blockers against 296e47d. Four are fixed — the Matching skip is gone from autoimpute.py, the is nnd_hotdeck_using_rpy2 identity check is gone, and the donation-class commit addresses the RANDwNND.hotdeck ordering. Thank you for turning those round so quickly.

The pickle blocker is only half fixed, and QRF still breaks. Your getattr guards cover dummy_processor, categorical_targets, boolean_targets and imputed_variables, but not the two attributes that caused the original failure:

sequential:        12 bare reads/writes, 0 guarded
_weighted_leaves:   6 bare reads/writes, 0 guarded

Reproduction on this head — fit, pickle, strip the new attributes to simulate a model pickled by an earlier release, predict:

ols: predict OK after stripping new attrs
qrf: RuntimeError: Failed to predict with QRF model:
     'QRFResults' object has no attribute 'sequential'

OLS is fixed. QRF fails at qrf.py:511, in QRFResults._predict:

if (
    quantiles is not None
    and self.sequential          # <- unguarded
    and len(self.imputed_variables) > 1
):

_weighted_leaves is reached next, in predict_quantiles_per_row.

This matters because policyengine-uk-data pickles fitted models to disk and reloads them (utils/qrf.py:78 dumps, :39 loads), with wealth.py, vat.py, income.py and services/etb.py all depending on it — so every cached UK model breaks on upgrade, at predict time inside a dataset build rather than at import.

Two lines would do it: getattr(self, "sequential", True) at :511 and getattr(self, "_weighted_leaves", None) at its read site. A __setstate__ on QRFResults and _QRFModel setting both defaults would be more durable, since it also covers attributes added later. Either way a .breaking.md fragment telling downstream to regenerate cached models is worth having, because a pickle written by 3.1.x and loaded by the next release will silently differ in behaviour even once it loads.

I have not pushed anything to this branch — you are clearly mid-flight on it, and I did not want to collide. Happy to push the two-line guard plus a round-trip regression test if that is useful.

The remaining blocker from my earlier review is the split, which is your call rather than a defect. Now that #214, #215, #216, #217 and #218 are all on main, the rebase should be much smaller than when I first suggested it.

This branch was successfully deployed

1 active deployment
Preview — 296e47dc Deployed Sep 21, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment