diff --git a/.plans/delivery-plan.md b/.plans/delivery-plan.md index 7d31163..e4f31b1 100644 --- a/.plans/delivery-plan.md +++ b/.plans/delivery-plan.md @@ -1291,13 +1291,16 @@ docs-only, +MCP and +skills. Both halves changed: the arms became `+skills` and `-no-skills` — the MCP server is read-only and ships in the CLI, so an MCP arm could only ever have moved investigate and resolve — and the suite grew. -What the schedule actually runs, as of 28 August: +What the schedule ran, 31 August to 28 September, before the benchmark came off it: | | Experiments | Pairs | Attempts | |---|---|---|---| | Weekly | 4 (frontier agents, plus the weak pair) | 4 x 19 = **76** | `runs=1` | | Monthly | 6 (adds the `-no-skills` twins) | 6 x 19 = **114** | `runs=1` | +Since 1 October the schedule runs the regression suite only, 3 x 6 = 18 pairs at two +attempts, and a benchmark matrix is dispatched against a bucket of work. + **The per-pair cost has not been re-measured since the suite grew**, and the two figures in this file disagree: the ten-scenario re-baseline above puts the mean at $1.44, while the most recent costing — a three-attempt matrix over nineteen scenarios @@ -1361,8 +1364,8 @@ start earning from day one. | Trigger | Runs | Why | |---|---|---| -| `schedule`, weekly | benchmark suite, frontier agents + weak pair | Published scores, and the weak pair is the only source of failures | -| `schedule`, monthly | benchmark suite, full matrix | Adds the `-no-skills` twins, whose delta does not move week to week | +| `schedule`, weekly | **regression suite, every experiment** | Catches a guarded mistake returning, and opens an issue when it does | +| `workflow_dispatch`, against a bucket of work | benchmark suite, the experiments that bucket needs | An eval change or a product change, measured by the run that follows it | | `repository_dispatch` from the docs repo, and PRs touching `skills/` | regression suite only | Cheap, catches hallucination regressions | | PR label `run-evals-changed` | only the scenarios changed in the PR | Scenario authoring loop | | `workflow_dispatch` | anything, by suite/eval/experiment | New model releases, one-offs | @@ -1373,8 +1376,9 @@ pair. Lift it, change the schedule block and the suite defaults, and add the sha loop from Phase 1 until the org key lands. **Where this stands today.** Pull requests run formatting, typecheck, unit tests and -build. `eval-refresh` runs on two schedules, weekly for the frontier agents plus the -weak pair and monthly for the full matrix, and on manual dispatch. Lifting the file +build. `eval-refresh` runs the regression suite weekly and nothing else on a schedule; +the benchmark is dispatched against a bucket of work. That changed on 1 October 2026 — +see Runs in AGENTS.md. Lifting the file wholesale turned out to carry more than the schedule: it arrived live on supabase/evals' nightly cron and fired against a half-built suite, and its matrix ran pairs concurrently, which one shared project cannot support. It is now diff --git a/.plans/evals-page-brief.md b/.plans/evals-page-brief.md index 676377c..8b4f867 100644 --- a/.plans/evals-page-brief.md +++ b/.plans/evals-page-brief.md @@ -32,11 +32,12 @@ several experiments, and one is capability-gated behind Outpost credentials so i deliberately reports a skip. "Not run" is a normal state here, not an edge case, and it has to read as distinct from a failure. -**Cells in the same grid are not the same age.** Frontier agents and the weak pair run -weekly; the `-no-skills` twins run monthly, because their measured delta is what got -cut to fit the budget. Half the columns can therefore be up to four weeks staler than -the other half. One "last updated" stamp for the table would be a false claim. -Freshness is per experiment. +**Cells in the same grid are not the same age, and nothing bounds the gap.** The +benchmark stopped running on a schedule on 1 October 2026: a run is dispatched against +a piece of work, so a column is as old as the last run that covered it. Until then the +frontier agents and the weak pair ran weekly and the `-no-skills` twins monthly, which +capped the gap at four weeks. There is no cap now. One "last updated" stamp for the +table would be a false claim; freshness is per experiment. ## The experiments, and what the pairing actually showed diff --git a/apps/framework/harness/run-eval.ts b/apps/framework/harness/run-eval.ts index 9333dde..a57fd95 100644 --- a/apps/framework/harness/run-eval.ts +++ b/apps/framework/harness/run-eval.ts @@ -834,7 +834,7 @@ async function main() { eval: ev.id, // When this run finished. Recorded per run rather than per // refresh, because the published grid is not one snapshot: the - // `-no-skills` twins refresh monthly and everything else weekly, + // a dispatched run covers only the experiments it asked for, // and a targeted re-run replaces some pairs and leaves the rest, // so two cells side by side can be a month apart. Without this // the page can only describe the schedule, which is not the same diff --git a/apps/framework/lib/provenance.ts b/apps/framework/lib/provenance.ts index 34c7873..7d61f95 100644 --- a/apps/framework/lib/provenance.ts +++ b/apps/framework/lib/provenance.ts @@ -12,8 +12,8 @@ import { * from a current one. * * `export-results --merge` carries forward any cell a run did not re-execute. - * That earns its place — the `-no-skills` twins refresh monthly and everything - * else weekly, so without it a weekly snapshot would have holes. What it lacks + * That earns its place — a dispatched run covers the experiments that run asked + * for, so without it a partial run's snapshot would have holes. What it lacks * is any notion of a row going *out of date*. On 25 August the published file * held 114 rows across six execution dates spanning thirteen days, 22 of them * from 13 August: measured before the sandbox CLI moved to 2.5.0, before fixed diff --git a/apps/framework/scripts/export-results.ts b/apps/framework/scripts/export-results.ts index c123eb2..bee69cc 100644 --- a/apps/framework/scripts/export-results.ts +++ b/apps/framework/scripts/export-results.ts @@ -75,8 +75,8 @@ const MERGE = rawArgs.includes('--merge'); * * Opt-in, because a snapshot with holes and a snapshot with silent thirteen-day- * old rows are both wrong and which is less wrong depends on what the snapshot - * is for. A release cut to report a measured change wants this on; a weekly - * refresh keeping the page populated probably does not. Without it the report + * is for. A release cut to report a measured change wants this on; a partial + * run keeping the page populated probably does not. Without it the report * below still prints, so the staleness is visible either way — which is the * actual defect in #60. Nothing was ever *said*. */ diff --git a/reference/design-tokens.md b/reference/design-tokens.md index 3586836..20a2cb2 100644 --- a/reference/design-tokens.md +++ b/reference/design-tokens.md @@ -127,7 +127,9 @@ because a flat delta honestly reported is worth more than an implied one, but do build the hierarchy around it. The spread that carries signal is model capability: the frontier agents pass nearly everything and a deliberately weaker model fails several. -**Cells in one grid are not the same age.** Frontier agents and the weak pair run -weekly; the `-no-skills` twins run monthly. A single "last updated" stamp for the table -would be wrong by up to four weeks on half the columns. Freshness is per experiment. -The design has no treatment for this because the cadence was set after it was drawn. +**Cells in one grid are not the same age, and since 1 October 2026 there is no bound +on how far apart they can be.** The benchmark came off the schedule, so cells are +measured when a run is dispatched against a piece of work. A single "last updated" +stamp for the table would be a false claim; freshness is per experiment. This used to +be "up to four weeks" — a weekly against a monthly — and that ceiling is gone, which +makes the gap worse rather than resolved. The design has no treatment for it.