Skip to content

[models] Keep the objectness scores aligned with the boxes when empty crops are dropped - #2164

Open
MohammadHijjawi97 wants to merge 1 commit into
mindee:mainfrom
MohammadHijjawi97:fix/objectness-scores-dropped-crops
Open

MohammadHijjawi97 wants to merge 1 commit into
mindee:mainfrom
MohammadHijjawi97:fix/objectness-scores-dropped-crops

Conversation

@MohammadHijjawi97

Copy link
Copy Markdown
Contributor

This PR:

  • Fixes the objectness scores of OCRPredictor / KIEPredictor getting shifted when a detected box yields an empty crop. _prepare_crops drops such boxes before recognition, but the objectness scores were detached from the boxes beforehand and kept their original length, so every word after a dropped box was built with the score of an earlier box (the dropped box's own score for the first word after it). _prepare_crops now filters the scores with the same mask as the boxes, and both predictors use the filtered scores.

    import numpy as np
    from torch import nn
    from doctr.file_utils import CLASS_NAME
    from doctr.models.predictor import OCRPredictor
    
    class FixedDet(nn.Module):  # middle box lies on the right page border -> empty crop
        def forward(self, pages, **kwargs):
            boxes = np.array([[0.1, 0.1, 0.3, 0.2, 0.9], [1.0, 0.5, 1.0, 0.6, 0.1], [0.5, 0.1, 0.7, 0.2, 0.8]], dtype=np.float32)
            return [{CLASS_NAME: boxes} for _ in pages]
    
    class DummyReco(nn.Module):
        def forward(self, crops, **kwargs):
            return [("word", 1.0) for _ in crops]
    
    page = OCRPredictor(FixedDet(), DummyReco())([np.full((100, 200, 3), 255, dtype=np.uint8)]).pages[0]
    print([(w.geometry[0][0], round(w.objectness_score, 2)) for w in page.blocks[0].lines[0].words])

    Before: [(0.1, 0.9), (0.5, 0.1)] (the word at x=0.5 carries the dropped box's score); after: [(0.1, 0.9), (0.5, 0.8)]. Same for KIEPredictor.

    Empty crops come from degenerate boxes (clipped onto the page border, sub-pixel wide, or produced by a custom detection model), so this is an edge case, but the scores were silently wrong when it happened.

  • Adds test_predictors_keep_objectness_scores_aligned (OCR and KIE), which fails without the fix.

Checks: pytest tests/pytorch/test_models_zoo_pt.py tests/common/test_models_builder.py tests/common/test_models.py (all tests that don't need to download pretrained weights / test assets pass; I ran offline), ruff check / ruff format --check, mypy doctr/.

Any feedback is welcome 🤗

… crops are dropped

`_prepare_crops` drops the boxes whose crop is empty before recognition, but
the objectness scores were detached earlier and kept their original length.
Every word after a dropped box was then built with the score of an earlier
box (the dropped one, for the first word after it), in both `OCRPredictor`
and `KIEPredictor`. The scores are now filtered with the same mask as the
boxes.
@felixdittrich92 felixdittrich92 self-assigned this Oct 7, 2026
@codecov

codecov Bot commented Oct 7, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 97.54%. Comparing base (b72ca69) to head (4fc8ac1).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@           Coverage Diff           @@
##             main    #2164   +/-   ##
=======================================
  Coverage   97.54%   97.54%           
=======================================
  Files         170      170           
  Lines       10658    10659    +1     
=======================================
+ Hits        10396    10397    +1     
  Misses        262      262           
Flag Coverage Δ
unittests 97.54% <100.00%> (+<0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@felixdittrich92 felixdittrich92 added this to the 1.2.0 milestone Oct 7, 2026
@felixdittrich92 felixdittrich92 added type: bug Something isn't working module: models Related to doctr.models ext: tests Related to tests folder topic: text detection Related to the task of text detection labels Oct 7, 2026

@felixdittrich92 felixdittrich92 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @MohammadHijjawi97 LGTM 👍

But one thing we should do here is moving the hook after "Crop images".
Let's include this in the PR.

        # Apply hooks to loc_preds if any
        # for hook in self.hooks:
        #     loc_preds = hook(loc_preds)

        # Crop images
        crops, loc_preds, objectness_scores = self._prepare_crops(
            pages,
            loc_preds,
            objectness_scores,
            assume_straight_pages=self.assume_straight_pages,
            assume_horizontal=self._page_orientation_disabled,
        )

        # Apply hooks to loc_preds if any
        for hook in self.hooks:
            loc_preds = hook(loc_preds)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ext: tests Related to tests folder module: models Related to doctr.models topic: text detection Related to the task of text detection type: bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants