Skip to content

Retire the last nine cadence claims, and say the gap got worse - #90

Merged
leggetter merged 1 commit into
mainfrom
cadence-claims-sweep
Oct 1, 2026
Merged

leggetter merged 1 commit into
mainfrom
cadence-claims-sweep

Conversation

@leggetter

Copy link
Copy Markdown
Collaborator

Follows the second review on #86, which listed nine statements still describing a weekly benchmark and a monthly matrix as current fact.

File Was
reference/design-tokens.md "Frontier agents and the weak pair run weekly; the -no-skills twins run monthly"
.plans/evals-page-brief.md Same claim, same reasoning
.plans/delivery-plan.md "What the schedule actually runs, as of 28 August"; the triggers table's two schedule rows; "eval-refresh runs on two schedules"
apps/framework/lib/provenance.ts "the -no-skills twins refresh monthly and everything else weekly"
apps/framework/harness/run-eval.ts Same
apps/framework/scripts/export-results.ts "a weekly refresh keeping the page populated"

Two of these are not tidying

design-tokens.md and the page brief used the cadence to bound a real problem: cells in one grid are different ages, and a weekly against a monthly capped the gap at four weeks. That ceiling is gone. A dispatched benchmark means a column is as old as the last run that covered it.

So both now say there is no bound, and that the problem is worse rather than resolved. Deleting the number quietly would have hidden a design gap that the page still has no treatment for.

The three source comments

They justified --merge and the per-row ranAt by the twins refreshing monthly. The justification survives the change for a different reason, which is what they now say: a dispatched run covers only the experiments it asked for, so a partial run's snapshot still has holes without the merge.

Kept deliberately

The plan's "what the schedule actually runs" table stays, relabelled as the record of 31 August to 28 September, because the cost measurements below it are measurements of that period and would be unreadable without it.

🤖 Generated with Claude Code

https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK

The second review on #86 found nine statements still describing a weekly
benchmark and a monthly matrix as current. Six were prose; three were source
comments justifying behaviour that still exists for a different reason.

Two of them matter beyond tidiness, and both are now honest rather than merely
updated. `reference/design-tokens.md` and the page brief said cells in one grid
can be "up to four weeks" apart, with a weekly against a monthly as the
ceiling. That ceiling is gone. A dispatched benchmark means a column is as old
as the last run that covered it, so the problem the design has no treatment for
is larger than it was, not resolved. Saying "no bound" is the point; quietly
deleting the number would have hidden it.

The three source comments justified `--merge` and `ranAt` by the twins
refreshing monthly. The justification survives the cadence: a dispatched run
covers the experiments that run asked for, so a partial run's snapshot still
has holes without the merge, and two cells side by side can still be far apart.

The plan's triggers table listed two benchmark crons; it now lists the
regression cron and dispatch-against-a-bucket. Its "what the schedule actually
runs" table is kept as the record of 31 August to 28 September, labelled as
such, because the costs below it are measurements of that period.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK
@leggetter
leggetter merged commit dc4055e into main Oct 1, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant