Prune failed Matching trials instead of scoring the training mean - #215
Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
|
Reviewed with a fresh pass. The change is correct and the new tests are the strongest part of it: Two claims in the description checked and confirmed: To doBefore merging
Worth doing
Optional
On #219: it contains this fix and goes further — it also prunes on silently incomplete results ( Caveat worth recording: the changed lines have not been exercised against real StatMatch anywhere I could run. |
|
Review of 626278c: The remaining merge blocker is the inherited Lint failure. One nonblocking result-metadata defect also needs correction. Merge blocker — incorporate #218 and rerun CI. The Lint job fails because Ruff would reformat five unchanged Markdown files: Nonblocking — attach Validation at the reviewed head: 71 affected tests passed; the optional real-R Matching module skipped locally. Deterministic backends exercised actual Optuna pruning, all-pruned studies, partial output, index preservation, counter resets, and missing-target-only counting. A separate public-API reproduction confirmed the metadata discrepancy above. Unit/smoke and changelog CI passed; Lint failed. Integration: read-only merge checks found no textual conflicts in the cumulative #218 + #214 + #215 + #216 + #217 tree. This is not a combined-suite result. #219 conflicts with #214, #215, #217, and #218, so its later rebase must retain the seed, Matching, API, and formatter fixes. The PR description cites an older final commit and historical green checks; update it to the resulting implementation and current validation. |
Both tuning handlers caught every exception and substituted the training mean, with no log and no counter, and that score went straight into the Optuna objective. A mean-predictor is not a neutral score - on a low-signal target it can beat a genuine matching fit on quantile loss - so a parameter set under which matching always failed could be selected as best and reported as the winning method. Both now log the exception and raise TrialPruned. The predict path keeps its NaN fill, which is the right behaviour there, but now reports the total number of unmatched records rather than leaving silent NaN blocks. Fixes #210
Per review, three follow-ups. n_failed_records kept the previous successful call's value when a prediction raised before reaching _process_matching_results, so a caller reading it after an exception got a stale number. It is now reset on entry to _predict, and documented - including that Matching runs single-threaded, so the attribute is safe in practice but would race under concurrent calls on one fitted object. The count is also mirrored onto result.attrs['n_failed_records'], so a caller does not have to reach into the fitted model for it. An all-pruned study reported only that nothing succeeded. It now carries the most recent underlying failure, which is what a user needs when matching fails structurally - a bad dtype or a missing R package - and the cause was otherwise reachable only through __cause__. Also drops the docs reformatting, which belongs to #218.
The previous commit de-indented 'raise optuna.TrialPruned() from e' out of 'except Exception as e', so 'e' was unbound and the trial failed with UnboundLocalError instead of pruning - which the matching failure tests caught.
…ing failure metadata
626278c to
62d03ce
Compare
Critical fixedUpdated the branch onto current Should-address fixedDefault Verification
Pushed head: GitHub CI: All seven current-head checks passed, including lint, changelog, Python 3.12 smoke tests, Python 3.14 tests with R, and Vercel preview. GitHub reports this pull request as mergeable. (run). |
Fixes #210.
Matching tuning prunes trials when donor matching raises an error, including errors in later chunks. Failed trials cannot win model selection through fallback training means. If no trial completes, tuning raises a
ValueErrorthat includes the last matching error when available.Chunked predictions retain successful rows and return missing values for unmatched rows. Each prediction resets
n_failed_recordsand counts records with missing target values. Every returned prediction frame now includesattrs["n_failed_records"], for both default predictions and explicit quantiles; earlier frames retain their own counts after later calls.Validation: the affected Matching manifest passes 12 tests, covering actual Optuna pruning, all-pruned studies, partial output, indexes, counter resets, and result metadata. The four metadata regressions first produced two failing default-prediction cases and two passing explicit-quantile controls. The optional R/rpy2 integration module skips locally; these local results do not validate the real R backend. Verification confirmed unchanged program and test contents after the base rebase. The branch includes #218's shared lint repair; Ruff 0.16.7 formatting and
git diff --checkpass.Checks for current head
62d03ce.