Skip to content

Normalization preprocessing re-fits on the receiver, erasing the level difference between datasets #202

Description

@vahid-ahmadi

autoimpute(..., preprocessing={"x": "normalize"}) standardises the donor and the receiver independently, each to its own mean and standard deviation. That erases the level difference between the two datasets — which is precisely the information imputation is supposed to carry across.

comparisons/autoimpute_helpers.py:162-173 calls preprocess_data a second time on imputing_data[predictors] and discards the returned transform parameters into _, despite the comment saying "apply same transformations".

Reproduction

Donor x1 ~ N(10, 2), receiver x1 ~ N(30, 8), true relationship y = 3 * x1:

with preprocessing={"x1": "normalize"}:   imputed mean y = 29.83
without preprocessing:                    imputed mean y = 90.08
true                                      = 90.05

The imputation is wrong by a factor of three, with no warning. Any analysis that enabled normalisation has silently mapped the receiver onto the donor's location.

Fix

Apply the donor's fitted parameters to the receiver rather than re-fitting:

transform_result = preprocess_data(donor_data[predictors], ...)
receiver_scaled = apply_normalization(imputing_data[predictors],
                                      transform_result["normalization"])

Severity

Critical. It is silent, it is available through the documented public entry point, and it produces plausible output.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions