Skip to content

fix(workflow-executor): page through the segment when one padded page holds only known records - #1931

Merged
Scra3 merged 6 commits into
feature/prd-1183-runtime-automation-pollerfrom
feature/prd-1362-automation-poller-records-waiting-on-a-person-grow-the
Sep 29, 2026
Merged

Scra3 merged 6 commits into
feature/prd-1183-runtime-automation-pollerfrom
feature/prd-1362-automation-poller-records-waiting-on-a-person-grow-the

Conversation

@Scra3

@Scra3 Scra3 commented Sep 28, 2026 •

Copy link
Copy Markdown
Member

fixes PRD-1362

Targets the integration branch of #1906.

Why

Records waiting on a person keep their doing row in the automated inbox, and nothing bounds how many there are. That is validated behaviour. Past 150 known records, the poller replaced pk not_in with one padded page of at most 500 records, then filtered known ids on its side. Past about 480 known records, that single page held nothing but known records: the automation silently found no new record while the fallback inbox kept filling.

Change

  • Paging. For a single-column key, the padded read now pages through the segment, 500 records at a time, sorted by that key (ascending). It stops at the first of three conditions:
    • maxConcurrentRuns candidates collected (never more are sent, the same ceiling as the not_in path);
    • a page shorter than requested, meaning the end of the segment;
    • MAX_PADDED_PAGES = 5 pages read.
  • Warn log. The padded candidate read found no new record within its page cap fires when the cap is reached with no candidate while the segment went on. It carries known, pagesRead and paddedPageReason.
  • Failure on a later page. A failure on page 2 or later keeps what the earlier pages found, with a warn. A failure before any candidate was found, on page 1 or later, still fails the read.
  • Composite keys keep the single unsorted padded page they had before. agent-client sorts on one field only, so offset pages over a composite key tied on its first column would overlap or skip.
  • Log field. candidatePagesRead is added to Automated inbox polled.
  • Port and adapter. The segment reader port gets pageNumber and sortByPrimaryKey, and agent-client already supports both.

Unchanged:

  • the not_in path and its capability check (PRD-1279), so an inbox with 150 known records or fewer still costs one read;
  • reconcileClosed;
  • the sync contract.

The server is not touched, and there is no contract change.

Load on the customer's agent

  • At most 5 reads of 500 ids, key only, per inbox per sweep, and only when the inbox knows more than 150 records.
  • That supports about 2,480 records waiting on a person per inbox. That is about 10 hours of a workflow broken for every record, at Qonto's measured volume.
  • Page size and cap were agreed with the owner.
  • A sort on the key is accepted by the Node agent (field existence check only) and by forest-rails (ORDER BY after the segment scope). Datasources whose key is not sortable ignore it.

Tests

  • Poller:
    • it reads pages 2 and 3 when page 1 holds only known records;
    • it stops once enough candidates are found;
    • it stops at the end of the segment without warning;
    • it keeps earlier candidates when a later page fails;
    • it fails the read when page 1 fails, or when a later page fails before any candidate;
    • it caps at 5 pages, then warns;
    • it sends at most maxConcurrentRuns candidates;
    • it takes from a later page only what the earlier pages left of that budget;
    • it does not spend that budget on a record an earlier page already brought;
    • it pages the same way when the fallback comes from a refused not_in;
    • it reads a composite key in one unsorted page.
  • Adapter: page[number], sort on the first key column, and no sort by default.
  • Every behaviour was mutation-checked. The package suite passes (1971 tests).

🤖 Generated with Claude Code

Note

[!NOTE]

Fix padded candidate reads to page through segments of known records

  • AutomationPoller.readPaddedCandidates now walks up to MAX_PADDED_PAGES (5) ordered pages when a padded page holds only known records, instead of stopping after one page
  • Paging applies only to single-column primary keys: those reads use ascending first-key ordering. Composite-key reads remain one unsorted page
  • ListSegmentRecordIdsQuery gained optional page-number and sort controls, and AgentClientSegmentReader.listRecordIds serializes them into agent requests
  • The scan stops when the inbox capacity is filled or the segment ends, and keeps candidates from earlier pages if a later page fails. A warning is emitted when the five-page cap is exhausted with no new candidate, and the automated-inbox poll log now includes the pages-read count
  • Risk: composite-key inboxes in automation-poller.ts still cannot page past a fully-known padded page; check readPaddedCandidates if that path needs coverage

Changes since #1931 opened

  • Added a Jest test case for AutomationPoller.cycle scheduling that verifies pagination budget handling across multiple pages [d092bf9]
  • Modified the candidate record ID accumulation loop in AutomationPoller to exclude IDs already present in the candidates set in addition to those in knownSet, preventing duplicate record IDs from being accumulated across multiple pages within the same sweep [657f62d]
  • Added test cases in automation-poller.test.ts to verify that candidate selection does not waste run budget on duplicate records appearing in subsequent pages and that paging behavior remains consistent when falling back due to a refused not_in filter [657f62d]

Macroscope summarized 9b7663a.

@Scra3 Scra3 self-assigned this Sep 28, 2026
@linear-code

linear-code Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

PRD-1362

PRD-1408

@qltysh

qltysh Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

3 new issues

Tool Category Rule Count
qlty Structure High total complexity (count = 84) 1
qlty Structure Function with many parameters (count = 5): readPaddedCandidates 1
qlty Structure Function with high complexity (count = 11): readPaddedCandidates 1

Comment thread packages/workflow-executor/src/automation-poller.ts Outdated
@qltysh

qltysh Bot commented Sep 28, 2026

Copy link
Copy Markdown

Qlty


Coverage Impact

This PR will not change total coverage.

Modified Files with Diff Coverage (2)

RatingFile% DiffUncovered Line #s
Coverage rating: A Coverage rating: A
packages/workflow-executor/src/automation-poller.ts100.0%
Coverage rating: A Coverage rating: A
.../workflow-executor/src/adapters/agent-client-segment-reader.ts100.0%
Total100.0%
🚦 See full report on Qlty Cloud »

🛟 Help
  • Diff Coverage: Coverage for added or modified lines of code (excludes deleted files). Learn more.

  • Total Coverage: Coverage for the whole repository, calculated as the sum of all File Coverage. Learn more.

  • File Coverage: Covered Lines divided by Covered Lines plus Missed Lines. (Excludes non-executable lines including blank lines and comments.)

    • Indirect Changes: Changes to File Coverage for files that were not modified in this PR. Learn more.

Comment thread packages/workflow-executor/src/automation-poller.ts Outdated
Comment thread packages/workflow-executor/src/automation-poller.ts

page
.filter(recordId => !knownSet.has(recordId))
.slice(0, config.maxConcurrentRuns - candidates.size)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Test gap: cross-page candidate-cap truncation is never exercised.

This .slice(0, config.maxConcurrentRuns - candidates.size) runs on every page, but across the padded-paging suite the subtraction only ever executes as 20 - 0 — the cap is either satisfied entirely from page 1, or the multi-page tests finish well under the cap (finals of 2 and 7). No test accumulates candidates across pages so that a later page must be truncated by a non-zero running total.

That's the exact budget this line enforces, and the logic the recent "cap padded candidates at the run budget" commit was circling. A mutation replacing config.maxConcurrentRuns - candidates.size with a constant config.maxConcurrentRuns (over-slicing every page instead of just the remaining budget) would pass the whole suite.

Suggested case: page 1 yields 15 fresh records (below the 20 cap, so the loop continues), page 2 yields 15 more → assert sync receives exactly 20 candidates and page 2 contributed only 5.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Valid, fixed in d092bf9. Added "should take from a later page only what the earlier pages left of the run budget": page 1 brings 15 fresh ids, page 2 brings 15 more, and sync must receive the 15 plus only the first 5 of page 2. Red with the mutation you describe (.slice(0, config.maxConcurrentRuns) sends 30).

Comment on lines +637 to +640
page
.filter(recordId => !knownSet.has(recordId))
.slice(0, config.maxConcurrentRuns - candidates.size)
.forEach(id => candidates.add(id));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The .slice() budget is maxConcurrentRuns - candidates.size, but the filter only removes ids in knownSet — not ids already collected in candidates from an earlier page. Reads are sequential, un-snapshotted offset pages (no cursor). If a record shifts across a page boundary between reads (a waiting-on-a-person record released, or a row inserted, between page 1 and page 2), a previously-seen id can reappear early in a later page: it passes the knownSet filter, consumes a slice slot, but candidates.add(id) is a no-op on the Set — so a genuinely new record past the cutoff on that page is dropped even though budget was still available.

Impact is bounded (the next iteration recomputes the budget, so a following page recovers the slack unless reachedEnd/MAX_PADDED_PAGES hits first — worst case one sweep under-fills the run budget), but it's a real gap and it's exactly the class of bug this PR guards against.

Suggested fix — also filter against what's already collected:

.filter(recordId => !knownSet.has(recordId) && !candidates.has(recordId))

Note: the Set dedup is never exercised today because every test uses disjoint page ranges — the same tests would pass with a plain array + push.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Valid, fixed in 657f62d. The filter now also drops ids already collected from an earlier page: !knownSet.has(recordId) && !candidates.has(recordId). Covered by "should not spend the run budget on a record an earlier page already brought": page 2 repeats fresh-14 from page 1, and sync must still receive the 15 of page 1 plus 5 genuinely new ids of page 2 (red without the filter, it sends 19).

);
});

describe('when one padded page holds nothing but known records', () => {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Coverage gap: the multi-page loop is never exercised via the not-in-refused route. readPaddedCandidates is now reached from two call sites — the direct paddedPageReason branch and the not_in-refused catch block (which also threads the new logContext parameter). Every test in this block uses too-many-known-records or the composite-key reason; the existing not-in-refused tests (~lines 704-774) return a short single page, so the loop never advances past page 1 on that path.

A bug specific to that route — wrong paddedPageReason in the cap-warning log, or knownSet computed differently when entered via the catch block — wouldn't be caught. Suggest adding one not-in-refused-flavored multi-page (or 5-page-cap) test.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Valid, added in 657f62d: "should page the same way when the fallback comes from a refused not_in". On that route the page size is maxConcurrentRuns + known, so a full page always holds a full budget and paging only kicks in once the 500 cap bites. The test uses maxConcurrentRuns: 480 with 30 known records: not_in is refused, page 1 brings 470 fresh ids, page 2 is read with pageNumber: 2 and sortByPrimaryKey: true, and the poll log carries paddedPageReason: "not-in-refused" and candidatePagesRead: 2.

Comment thread packages/workflow-executor/src/automation-poller.ts
Comment thread packages/workflow-executor/src/ports/segment-reader-port.ts
@Scra3
Scra3 force-pushed the feature/prd-1183-runtime-automation-poller branch from 03a2448 to a87e99d Compare September 29, 2026 09:57

@ShohanRahman ShohanRahman left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both raised points (observability of a partial padded read on a persistent later-page failure, and the port not encoding the single-column-key paging invariant) were answered convincingly:

  • The partial read is already machine-distinguishable: a distinct A later padded candidate page failed… Warn carrying inboxId/pagesRead/error once per sweep, rate-alertable, with Error correctly reserved for the no-progress case. Records aren't lost — re-read on the next sweep. Adding per-inbox failure state just for a log level isn't worth it.
  • The paging invariant lives in the sole caller (gated on a single-column key, sending sortByPrimaryKey: pageable) and is pinned by the composite-key test. A sortByPrimaryKey: true type wouldn't cover the composite case anyway. Fair to fold the fields together only when a second caller appears.

Core paging/budget/end-of-segment/error-handling logic reviewed as correct and well-covered. LGTM 👍

alban bertolini and others added 6 commits September 29, 2026 11:59
… holds only known records

Records waiting on a person stay known for as long as nobody handles them.
Past about 480 of them, the single padded page held none but them and the
automation silently found nothing. The padded read now pages, sorted by the
first key column, up to five pages of 500, and warns when that finds none.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…a later page fails

Review follow-up: a timeout on page 2 no longer sends an empty sync, and
two comments no longer overclaim.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ded a candidate

A later page failing before any candidate no longer reads as a page cap
reached with nothing new.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… composite keys on one page

Review follow-ups (Macroscope): a 500-record page no longer sends 500
candidates, and a composite key, which agent-client can only sort on its
first column, keeps the single unsorted page offset paging cannot walk.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A later page must only add what the earlier pages left of maxConcurrentRuns.
No test accumulated candidates across pages, so slicing every page at the
full budget passed the suite.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…rlier padded page brought

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@Scra3
Scra3 force-pushed the feature/prd-1362-automation-poller-records-waiting-on-a-person-grow-the branch from 657f62d to 9b7663a Compare September 29, 2026 10:04
@Scra3
Scra3 merged commit f53dfa9 into feature/prd-1183-runtime-automation-poller Sep 29, 2026
37 checks passed
@Scra3
Scra3 deleted the feature/prd-1362-automation-poller-records-waiting-on-a-person-grow-the branch September 29, 2026 10:12
Scra3 added a commit that referenced this pull request Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants