Skip to content

Add tximport-style transcript-length normalization and pytximport support - #451

Open
gitbenlewis wants to merge 10 commits into
scverse:mainfrom
gitbenlewis:main
Open

gitbenlewis wants to merge 10 commits into
scverse:mainfrom
gitbenlewis:main

Conversation

@gitbenlewis

Copy link
Copy Markdown

Reference Issue or PRs

Closes #305.
Closes #359.
Related to #412.

What does your PR implement? Be specific.

This PR adds DESeq2-style transcript-length
normalization and direct compatibility
with AnnData objects produced by
[pytximport](https://github.com/complextissue/
pytximport).

The generic API remains available:

dds = DeseqDataSet(
    counts=estimated_counts,
    metadata=metadata,
    transcript_lengths=average_transcript_lengths,
    design="~condition",
)

Compatible pytximport objects can also be passed
directly:

dds = DeseqDataSet(adata=txi, design="~condition")

The implementation:

  • Rounds unscaled estimated counts, matching
    DESeq2.
  • Stores average transcript lengths in
    layers["avg_tx_length"].
  • Constructs sample-by-gene normalization factors
    containing both transcript-length
    and relative library-size components.
  • Stores the final factors in
    layers["normalization_factors"].
  • Uses the factor matrix throughout dispersion
    fitting, LFC fitting, Wald tests,
    shrinkage, Cook's outlier replacement/refitting,
    and VST.
  • Supports the ratio and poscounts size-factor
    methods.
  • Preserves sparse CSR and CSC counts through
    inference boundaries.
  • Does not add pytximport as a runtime
    dependency.

Length-source precedence is deterministic:

  1. Explicit transcript_lengths
  2. adata.layers["avg_tx_length"]
  3. Compatible pytximport fields

For pytximport input, only unscaled estimated
counts with
adata.uns["counts_from_abundance"] is None are
accepted. Abundance-scaled modes
are rejected to prevent applying transcript-length
correction twice.

Validation

Local verification completed successfully:

  • Full test suite: 104 passed
  • Dense-versus-CSR/CSC parity through dispersion
    fitting, LFCs, Wald tests,
    shrinkage, VST, Cook's distances, replacement,
    and refitting
  • AnnData 0.11.4 compatibility checks passed
  • Ruff, formatting, mypy, pre-commit, and Sphinx
    warnings-as-errors passed

Current limitations

  • Direct pytximport normalization requires an in-
    memory AnnData object.
  • Pytximport's current obsm["length"] matrix has
    no gene-axis labels, so it must
    remain synchronized when genes are subset or
    reordered.
  • Iterative size-factor fitting is not supported
    with transcript-length offsets.
  • VST transformation of an external count matrix
    with separate transcript lengths
    is not currently supported.

Related work and acknowledgements

This implementation began independently with the
generic transcript_lengths
API and was extended with direct pytximport
compatibility following discussion in
#412.

Thank you to @maltekuehl for the earlier draft
implementation and design work in
#412, and to @Zethson for encouraging this work to
be taken forward. The original
tximport, DESeq2, and pytximport work is credited
in the documentation.

@gitbenlewis

gitbenlewis commented Jul 29, 2026 •

Copy link
Copy Markdown
Author

Hi @BorisMuzellec — just checking in on PR #451 when you have a chance. All current checks are passing. Could you share an update on the review, or let me know if there’s anything else you’d like me to address? Thanks!
@grst

@sriramdevadas

Copy link
Copy Markdown

@Zethson, would it be useful if I verified this against R's DESeqDataSetFromTximport on a public salmon dataset and contributed the R reference outputs as fixtures in the repo's existing test format? Glad to do it, or to stay out of the way if you have it covered. (cc @gitbenlewis)

@gitbenlewis

gitbenlewis commented Sep 21, 2026 •

Copy link
Copy Markdown
Author

@Zethson, would it be useful if I verified this against R's DESeqDataSetFromTximport on a public salmon dataset and contributed the R reference outputs as fixtures in the repo's existing test format? Glad to do it, or to stay out of the way if you have it covered. (cc @gitbenlewis)

Thanks! I have a separate PyDESeq2–R DESeq2 parity-testing repository, where I’ve added real-Salmon coverage using six public GEUVADIS samples imported independently through R tximport and pytximport. Import and normalization checks pass, and both PyDESeq2 input interfaces agree.
There are some remaining numerical parity differences that appear unrelated to this PR’s import and normalization changes, with details in the diagnosis document. I’ve started addressing these on a separate dispersion-parity branch. It contains partial improvements; full parity is not yet achieved. I suggest keeping that work in a separate PR.
The R reference outputs are now included as fixtures in the external repository. An independent check and help integrating them into PyDESeq2’s existing test format would still be very welcome.

@Zethson merge conflicts resolved : )

@codecov-commenter

codecov-commenter commented Sep 21, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 89.63%. Comparing base (95ace12) to head (9b42a30).

Additional details and impacted files
@@            Coverage Diff             @@
##             main     #451      +/-   ##
==========================================
+ Coverage   86.05%   89.63%   +3.58%     
==========================================
  Files          15       15              
  Lines        1276     1486     +210     
==========================================
+ Hits         1098     1332     +234     
+ Misses        178      154      -24     
Files with missing lines Coverage Δ
src/pydeseq2/dds.py 92.42% <100.00%> (+4.79%) ⬆️
src/pydeseq2/default_inference.py 97.82% <100.00%> (+1.72%) ⬆️
src/pydeseq2/dispersions.py 96.93% <100.00%> (+0.09%) ⬆️
src/pydeseq2/ds.py 90.14% <100.00%> (+0.14%) ⬆️
src/pydeseq2/inference.py 100.00% <100.00%> (ø)
src/pydeseq2/misc.py 90.90% <100.00%> (+4.24%) ⬆️

... and 1 file with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Plans to add DESeqDataSetFromTximport? Add support for sample-/gene-dependent normalization factors (e.g., length offsets from pytximport)

3 participants