Add tximport-style transcript-length normalization and pytximport support - #451
gitbenlewis wants to merge 10 commits into
Conversation
|
Hi @BorisMuzellec — just checking in on PR #451 when you have a chance. All current checks are passing. Could you share an update on the review, or let me know if there’s anything else you’d like me to address? Thanks! |
|
@Zethson, would it be useful if I verified this against R's DESeqDataSetFromTximport on a public salmon dataset and contributed the R reference outputs as fixtures in the repo's existing test format? Glad to do it, or to stay out of the way if you have it covered. (cc @gitbenlewis) |
Thanks! I have a separate PyDESeq2–R DESeq2 parity-testing repository, where I’ve added real-Salmon coverage using six public GEUVADIS samples imported independently through R tximport and pytximport. Import and normalization checks pass, and both PyDESeq2 input interfaces agree. @Zethson merge conflicts resolved : ) |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #451 +/- ##
==========================================
+ Coverage 86.05% 89.63% +3.58%
==========================================
Files 15 15
Lines 1276 1486 +210
==========================================
+ Hits 1098 1332 +234
+ Misses 178 154 -24
🚀 New features to boost your workflow:
|
Reference Issue or PRs
Closes #305.
Closes #359.
Related to #412.
What does your PR implement? Be specific.
This PR adds DESeq2-style transcript-length
normalization and direct compatibility
with AnnData objects produced by
[
pytximport](https://github.com/complextissue/pytximport).
The generic API remains available:
Compatible pytximport objects can also be passed
directly:
The implementation:
DESeq2.
layers["avg_tx_length"].containing both transcript-length
and relative library-size components.
layers["normalization_factors"].fitting, LFC fitting, Wald tests,
shrinkage, Cook's outlier replacement/refitting,
and VST.
ratioandposcountssize-factormethods.
inference boundaries.
pytximportas a runtimedependency.
Length-source precedence is deterministic:
transcript_lengthsadata.layers["avg_tx_length"]For pytximport input, only unscaled estimated
counts with
adata.uns["counts_from_abundance"] is Noneareaccepted. Abundance-scaled modes
are rejected to prevent applying transcript-length
correction twice.
Validation
Local verification completed successfully:
fitting, LFCs, Wald tests,
shrinkage, VST, Cook's distances, replacement,
and refitting
warnings-as-errors passed
Current limitations
memory AnnData object.
obsm["length"]matrix hasno gene-axis labels, so it must
remain synchronized when genes are subset or
reordered.
with transcript-length offsets.
with separate transcript lengths
is not currently supported.
Related work and acknowledgements
This implementation began independently with the
generic
transcript_lengthsAPI and was extended with direct pytximport
compatibility following discussion in
#412.
Thank you to @maltekuehl for the earlier draft
implementation and design work in
#412, and to @Zethson for encouraging this work to
be taken forward. The original
tximport, DESeq2, and pytximport work is credited
in the documentation.