From 14e73482ca0bdb9732215f511413251a2fa31c78 Mon Sep 17 00:00:00 2001 From: Phil Leggetter Date: Thu, 1 Oct 2026 14:38:24 +0100 Subject: [PATCH] Retire the last nine cadence claims, and say the gap got worse The second review on #86 found nine statements still describing a weekly benchmark and a monthly matrix as current. Six were prose; three were source comments justifying behaviour that still exists for a different reason. Two of them matter beyond tidiness, and both are now honest rather than merely updated. `reference/design-tokens.md` and the page brief said cells in one grid can be "up to four weeks" apart, with a weekly against a monthly as the ceiling. That ceiling is gone. A dispatched benchmark means a column is as old as the last run that covered it, so the problem the design has no treatment for is larger than it was, not resolved. Saying "no bound" is the point; quietly deleting the number would have hidden it. The three source comments justified `--merge` and `ranAt` by the twins refreshing monthly. The justification survives the cadence: a dispatched run covers the experiments that run asked for, so a partial run's snapshot still has holes without the merge, and two cells side by side can still be far apart. The plan's triggers table listed two benchmark crons; it now lists the regression cron and dispatch-against-a-bucket. Its "what the schedule actually runs" table is kept as the record of 31 August to 28 September, labelled as such, because the costs below it are measurements of that period. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK --- .plans/delivery-plan.md | 14 +++++++++----- .plans/evals-page-brief.md | 11 ++++++----- apps/framework/harness/run-eval.ts | 2 +- apps/framework/lib/provenance.ts | 4 ++-- apps/framework/scripts/export-results.ts | 4 ++-- reference/design-tokens.md | 10 ++++++---- 6 files changed, 26 insertions(+), 19 deletions(-) diff --git a/.plans/delivery-plan.md b/.plans/delivery-plan.md index 7d31163..e4f31b1 100644 --- a/.plans/delivery-plan.md +++ b/.plans/delivery-plan.md @@ -1291,13 +1291,16 @@ docs-only, +MCP and +skills. Both halves changed: the arms became `+skills` and `-no-skills` — the MCP server is read-only and ships in the CLI, so an MCP arm could only ever have moved investigate and resolve — and the suite grew. -What the schedule actually runs, as of 28 August: +What the schedule ran, 31 August to 28 September, before the benchmark came off it: | | Experiments | Pairs | Attempts | |---|---|---|---| | Weekly | 4 (frontier agents, plus the weak pair) | 4 x 19 = **76** | `runs=1` | | Monthly | 6 (adds the `-no-skills` twins) | 6 x 19 = **114** | `runs=1` | +Since 1 October the schedule runs the regression suite only, 3 x 6 = 18 pairs at two +attempts, and a benchmark matrix is dispatched against a bucket of work. + **The per-pair cost has not been re-measured since the suite grew**, and the two figures in this file disagree: the ten-scenario re-baseline above puts the mean at $1.44, while the most recent costing — a three-attempt matrix over nineteen scenarios @@ -1361,8 +1364,8 @@ start earning from day one. | Trigger | Runs | Why | |---|---|---| -| `schedule`, weekly | benchmark suite, frontier agents + weak pair | Published scores, and the weak pair is the only source of failures | -| `schedule`, monthly | benchmark suite, full matrix | Adds the `-no-skills` twins, whose delta does not move week to week | +| `schedule`, weekly | **regression suite, every experiment** | Catches a guarded mistake returning, and opens an issue when it does | +| `workflow_dispatch`, against a bucket of work | benchmark suite, the experiments that bucket needs | An eval change or a product change, measured by the run that follows it | | `repository_dispatch` from the docs repo, and PRs touching `skills/` | regression suite only | Cheap, catches hallucination regressions | | PR label `run-evals-changed` | only the scenarios changed in the PR | Scenario authoring loop | | `workflow_dispatch` | anything, by suite/eval/experiment | New model releases, one-offs | @@ -1373,8 +1376,9 @@ pair. Lift it, change the schedule block and the suite defaults, and add the sha loop from Phase 1 until the org key lands. **Where this stands today.** Pull requests run formatting, typecheck, unit tests and -build. `eval-refresh` runs on two schedules, weekly for the frontier agents plus the -weak pair and monthly for the full matrix, and on manual dispatch. Lifting the file +build. `eval-refresh` runs the regression suite weekly and nothing else on a schedule; +the benchmark is dispatched against a bucket of work. That changed on 1 October 2026 — +see Runs in AGENTS.md. Lifting the file wholesale turned out to carry more than the schedule: it arrived live on supabase/evals' nightly cron and fired against a half-built suite, and its matrix ran pairs concurrently, which one shared project cannot support. It is now diff --git a/.plans/evals-page-brief.md b/.plans/evals-page-brief.md index 676377c..8b4f867 100644 --- a/.plans/evals-page-brief.md +++ b/.plans/evals-page-brief.md @@ -32,11 +32,12 @@ several experiments, and one is capability-gated behind Outpost credentials so i deliberately reports a skip. "Not run" is a normal state here, not an edge case, and it has to read as distinct from a failure. -**Cells in the same grid are not the same age.** Frontier agents and the weak pair run -weekly; the `-no-skills` twins run monthly, because their measured delta is what got -cut to fit the budget. Half the columns can therefore be up to four weeks staler than -the other half. One "last updated" stamp for the table would be a false claim. -Freshness is per experiment. +**Cells in the same grid are not the same age, and nothing bounds the gap.** The +benchmark stopped running on a schedule on 1 October 2026: a run is dispatched against +a piece of work, so a column is as old as the last run that covered it. Until then the +frontier agents and the weak pair ran weekly and the `-no-skills` twins monthly, which +capped the gap at four weeks. There is no cap now. One "last updated" stamp for the +table would be a false claim; freshness is per experiment. ## The experiments, and what the pairing actually showed diff --git a/apps/framework/harness/run-eval.ts b/apps/framework/harness/run-eval.ts index 9333dde..a57fd95 100644 --- a/apps/framework/harness/run-eval.ts +++ b/apps/framework/harness/run-eval.ts @@ -834,7 +834,7 @@ async function main() { eval: ev.id, // When this run finished. Recorded per run rather than per // refresh, because the published grid is not one snapshot: the - // `-no-skills` twins refresh monthly and everything else weekly, + // a dispatched run covers only the experiments it asked for, // and a targeted re-run replaces some pairs and leaves the rest, // so two cells side by side can be a month apart. Without this // the page can only describe the schedule, which is not the same diff --git a/apps/framework/lib/provenance.ts b/apps/framework/lib/provenance.ts index 34c7873..7d61f95 100644 --- a/apps/framework/lib/provenance.ts +++ b/apps/framework/lib/provenance.ts @@ -12,8 +12,8 @@ import { * from a current one. * * `export-results --merge` carries forward any cell a run did not re-execute. - * That earns its place — the `-no-skills` twins refresh monthly and everything - * else weekly, so without it a weekly snapshot would have holes. What it lacks + * That earns its place — a dispatched run covers the experiments that run asked + * for, so without it a partial run's snapshot would have holes. What it lacks * is any notion of a row going *out of date*. On 25 August the published file * held 114 rows across six execution dates spanning thirteen days, 22 of them * from 13 August: measured before the sandbox CLI moved to 2.5.0, before fixed diff --git a/apps/framework/scripts/export-results.ts b/apps/framework/scripts/export-results.ts index c123eb2..bee69cc 100644 --- a/apps/framework/scripts/export-results.ts +++ b/apps/framework/scripts/export-results.ts @@ -75,8 +75,8 @@ const MERGE = rawArgs.includes('--merge'); * * Opt-in, because a snapshot with holes and a snapshot with silent thirteen-day- * old rows are both wrong and which is less wrong depends on what the snapshot - * is for. A release cut to report a measured change wants this on; a weekly - * refresh keeping the page populated probably does not. Without it the report + * is for. A release cut to report a measured change wants this on; a partial + * run keeping the page populated probably does not. Without it the report * below still prints, so the staleness is visible either way — which is the * actual defect in #60. Nothing was ever *said*. */ diff --git a/reference/design-tokens.md b/reference/design-tokens.md index 3586836..20a2cb2 100644 --- a/reference/design-tokens.md +++ b/reference/design-tokens.md @@ -127,7 +127,9 @@ because a flat delta honestly reported is worth more than an implied one, but do build the hierarchy around it. The spread that carries signal is model capability: the frontier agents pass nearly everything and a deliberately weaker model fails several. -**Cells in one grid are not the same age.** Frontier agents and the weak pair run -weekly; the `-no-skills` twins run monthly. A single "last updated" stamp for the table -would be wrong by up to four weeks on half the columns. Freshness is per experiment. -The design has no treatment for this because the cadence was set after it was drawn. +**Cells in one grid are not the same age, and since 1 October 2026 there is no bound +on how far apart they can be.** The benchmark came off the schedule, so cells are +measured when a run is dispatched against a piece of work. A single "last updated" +stamp for the table would be a false claim; freshness is per experiment. This used to +be "up to four weeks" — a weekly against a monthly — and that ceiling is gone, which +makes the gap worse rather than resolved. The design has no treatment for it.