fix(approvals): drain shielded outcome writes before shutdown, per gate, logged once (BACKLOG #2087) - #1814
Merged
Conversation
added 2 commits
September 29, 2026 15:53
…te, logged once (BACKLOG #2087) The approval gate's shielded outcome writes (BACKLOG #1562) had three gaps. - Nothing drained them before engine.stop() closed the store. ApprovalGate.drain() now waits, bounded by DRAIN_TIMEOUT_SECONDS, and the create_managed_app lifespan calls it before engine.stop(), guarded so a failure or the deadline never skips the stop. Writes still running at the deadline are logged with their approval ids. - The task set was module-global. It is now held per ApprovalGate instance. - A write that raised after its caller was cancelled was logged twice: once by the gate and once by asyncio.shield itself ("exception in shielded future", through the loop exception handler). The claim path had the same shape. Both now wait with asyncio.wait, which never cancels the task and reports nothing, so the gate's line is the only one. Folded in PR 1636 item 5: a resolve whose caller was cancelled and that then lost the race logs its 409 at WARNING without a traceback, not at ERROR. Limb 4 (a row left 'executing' after a failed status write) is NOT built. It waits on #1562's startup-reconciler and process-ownership design; test_a_failed_approved_write_still_writes_the_audit_row still pins today's behaviour. Tests: single-report caplog tests across the gate and asyncio loggers for a failed cancelled claim and for a cancelled resolve that loses the race; a cancelled claim that loses the race; drain wait, deadline and per-gate scope; the lifespan drain with a neutralised-drain control; a failing drain still reaches engine.stop(). BACKLOG #2087 Proposed PR title: fix(approvals): drain shielded outcome writes before shutdown, per gate, logged once (BACKLOG #2087) Proposed ledger banner: PARTIAL -- limbs 1, 2, 3 and 5 shipped (bounded per-gate drain before engine.stop(), one log per failed shielded write, PR 1636 item 5's lost-race 409 at WARNING, the two missing tests). Limb 4, a way out of a stuck 'executing' row, remains open on #1562's startup-reconciler design.
…ogger (BACKLOG #2087) From the code-review subagent at xhigh: - DRAIN_TIMEOUT_SECONDS drops from 10 to 3. It shares NSSM's 15 s graceful-stop window with the upload runner's 5 s stop, and engine.stop() still has to run. - drain() re-reads the task set until it is empty or the deadline passes, so a write started while it waits is drained too. New test pins it. - Only a 409 ApprovalError from an orphaned write logs at WARNING; any other ApprovalError keeps ERROR and its traceback. - Tests: collect garbage before counting reports, so an error nobody read cannot hide as one report; the no-drain control holds its write until the lifespan has exited instead of racing a fixed sleep against teardown. BACKLOG #2087 Proposed PR title: fix(approvals): drain shielded outcome writes before shutdown, per gate, logged once (BACKLOG #2087) Proposed ledger banner: PARTIAL -- limbs 1, 2, 3 and 5 shipped (bounded per-gate drain before engine.stop(), one log per failed shielded write, PR 1636 item 5's lost-race 409 at WARNING, the two missing tests). Limb 4, a way out of a stuck 'executing' row, remains open on #1562's startup-reconciler design.
Collaborator
Author
|
QA: code-review xhigh subagent, 10 findings, 5 fixed, 5 left (3: mid-drain cancel skips engine.stop, same shape as the existing flush guard, separate change; 5: unread claim error only on loop-exit cancel, settle is drained; 8: introspection link cosmetic; 9: embedded create_app path documented to call drain(); 10: dead guard removed in the #2 rewrite) |
wshallwshall
enabled auto-merge
September 29, 2026 21:27
added 3 commits
September 29, 2026 17:40
…g the installed handler (BACKLOG #2087) The two log-once tests asserted the running loop had no exception handler, so a shield's report would reach the 'asyncio' logger. The tests share one session loop, and any create_managed_app lifespan earlier on it installs the engine's last-resort handler (messagefoundry/last_resort.py, install_loop_exception_handler, called from the lifespan in messagefoundry/api/app.py) and never removes it. On the windows-2025 leg that ordering held, and the assertion failed. Each test now installs its own capturing handler for its duration, restores the previous one, and asserts that nothing reached it. Measured both ways: the new tests pass with the last-resort handler leaked in first and on a clean loop, and against the pre-fix approvals.py they fail in both states. BACKLOG #2087 Proposed PR title: fix(approvals): drain shielded outcome writes before shutdown, per gate, logged once (BACKLOG #2087) Proposed ledger banner: PARTIAL -- limbs 1, 2, 3 and 5 shipped (bounded per-gate drain before engine.stop(), one log per failed shielded write, PR 1636 item 5's lost-race 409 at WARNING, the two missing tests). Limb 4, a way out of a stuck 'executing' row, remains open on #1562's startup-reconciler design.
…s (BACKLOG #2087) Review repair. The capturing handler counted every loop exception, so garbage left by earlier tests on the shared session loop, collected inside the window, could fail a log-once test for a reason unrelated to it. _loop_reports() now collects before it installs its handler, and a failure names the exception. BACKLOG #2087 Proposed PR title: fix(approvals): drain shielded outcome writes before shutdown, per gate, logged once (BACKLOG #2087) Proposed ledger banner: PARTIAL -- limbs 1, 2, 3 and 5 shipped (bounded per-gate drain before engine.stop(), one log per failed shielded write, PR 1636 item 5's lost-race 409 at WARNING, the two missing tests). Limb 4, a way out of a stuck 'executing' row, remains open on #1562's startup-reconciler design.
Collaborator
Author
|
Lander review of the repair commits (
Low notes, none blocking:
Re-arming. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Batch 180 (webconsole), one item. Its own PR because it changes a dual-control security control.
BACKLOG #2087 -- approvals shielded-outcome follow-ups (limbs 1, 2, 3, 5 and PR 1636 item 5)
b180/2087-approvals-shield7d43dcde6577700555a66484f40cbc5648c964a87412632003af0d09a966f2460cdfd11e6819ab1aWhat changed (messagefoundry/api/approvals.py, the create_managed_app teardown in app.py, tests/test_approvals.py):
ApprovalGate.drain(timeout=3.0)waits for outcome writes still running, re-reading the set so a write started mid-drain is drained too. It cancels nothing, and logs any still running at the deadline at ERROR. The managed lifespan calls it before the summary-audit flush and beforeengine.stop(), guarded so neither a failure nor the deadline skipsengine.stop(). The gate is read throughgetattrbecause an early startup failure may not have built one. 3 s, not 10: it shares NSSM's 15 s graceful-stop window with the upload runner's 5 s stop._SHIELDEDset is gone;_shieldedis a gate method trackingself._inflight.asyncio.shielduses are replaced with_outlive_caller(asyncio.waitthentask.result()). A shield reports a late failure itself ("exception in shielded future"), so one failure was logged twice. The Builder measured this red on origin/main: both log-once tests saw exactly 2 reports. Note the readings missed the directasyncio.shield(claim)inapprove(); the row's "a failed settle logs twice" was accurate for that path.engine.stop(), order checked.Not built, fenced. Limb 4 (a row left
executingafter a failed status write) waits on #1562's startup-reconciler / process-ownership design. No route out ofexecutingwas written, andtest_a_failed_approved_write_still_writes_the_audit_rowis unchanged.Review notes left open (code-review xhigh, 10 findings, 5 fixed): a lifespan cancelled mid-drain skips
engine.stop()(the existing flush guard has the same shape; fixing both is a separate change); the embeddedcreate_app(engine=...)path never drains (the docstring tells embedders to calldrain()).Checks the Builder ran.
ruff check,ruff format --check,mypy messagefoundry(302 files),mypy --explicit-package-bases tests(1020 files) clean. pytest over approvals, summary flush, approver provenance, requester recheck, min dwell, apiclient approval hold, lifespan startup unwinds, audit integrity, logging, api_tls and neighbours: 487 + 151 + 59 passed before repairs, 122 + 10 passed after. Full suite not run locally.Legs not seen locally. The Postgres and SQL Server store-contract legs (no store method changed), and windows-service-smoke, which runs this teardown under NSSM.
Proposed ledger banner: PARTIAL -- limbs 1, 2, 3 and 5 and PR 1636 item 5 shipped (per-gate drain before engine.stop(), log once). Limb 4 remains open on #1562's startup-reconciler design.