Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 9 additions & 5 deletions .plans/delivery-plan.md
Original file line number Diff line number Diff line change
Expand Up @@ -1291,13 +1291,16 @@ docs-only, +MCP and +skills. Both halves changed: the arms became `+skills` and
`-no-skills` — the MCP server is read-only and ships in the CLI, so an MCP arm could
only ever have moved investigate and resolve — and the suite grew.

What the schedule actually runs, as of 28 August:
What the schedule ran, 31 August to 28 September, before the benchmark came off it:

| | Experiments | Pairs | Attempts |
|---|---|---|---|
| Weekly | 4 (frontier agents, plus the weak pair) | 4 x 19 = **76** | `runs=1` |
| Monthly | 6 (adds the `-no-skills` twins) | 6 x 19 = **114** | `runs=1` |

Since 1 October the schedule runs the regression suite only, 3 x 6 = 18 pairs at two
attempts, and a benchmark matrix is dispatched against a bucket of work.

**The per-pair cost has not been re-measured since the suite grew**, and the two
figures in this file disagree: the ten-scenario re-baseline above puts the mean at
$1.44, while the most recent costing — a three-attempt matrix over nineteen scenarios
Expand Down Expand Up @@ -1361,8 +1364,8 @@ start earning from day one.

| Trigger | Runs | Why |
|---|---|---|
| `schedule`, weekly | benchmark suite, frontier agents + weak pair | Published scores, and the weak pair is the only source of failures |
| `schedule`, monthly | benchmark suite, full matrix | Adds the `-no-skills` twins, whose delta does not move week to week |
| `schedule`, weekly | **regression suite, every experiment** | Catches a guarded mistake returning, and opens an issue when it does |
| `workflow_dispatch`, against a bucket of work | benchmark suite, the experiments that bucket needs | An eval change or a product change, measured by the run that follows it |
| `repository_dispatch` from the docs repo, and PRs touching `skills/` | regression suite only | Cheap, catches hallucination regressions |
| PR label `run-evals-changed` | only the scenarios changed in the PR | Scenario authoring loop |
| `workflow_dispatch` | anything, by suite/eval/experiment | New model releases, one-offs |
Expand All @@ -1373,8 +1376,9 @@ pair. Lift it, change the schedule block and the suite defaults, and add the sha
loop from Phase 1 until the org key lands.

**Where this stands today.** Pull requests run formatting, typecheck, unit tests and
build. `eval-refresh` runs on two schedules, weekly for the frontier agents plus the
weak pair and monthly for the full matrix, and on manual dispatch. Lifting the file
build. `eval-refresh` runs the regression suite weekly and nothing else on a schedule;
the benchmark is dispatched against a bucket of work. That changed on 1 October 2026 —
see Runs in AGENTS.md. Lifting the file
wholesale turned out to carry more than the schedule: it arrived live on
supabase/evals' nightly cron and fired against a half-built suite, and its matrix ran
pairs concurrently, which one shared project cannot support. It is now
Expand Down
11 changes: 6 additions & 5 deletions .plans/evals-page-brief.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,11 +32,12 @@ several experiments, and one is capability-gated behind Outpost credentials so i
deliberately reports a skip. "Not run" is a normal state here, not an edge case, and
it has to read as distinct from a failure.

**Cells in the same grid are not the same age.** Frontier agents and the weak pair run
weekly; the `-no-skills` twins run monthly, because their measured delta is what got
cut to fit the budget. Half the columns can therefore be up to four weeks staler than
the other half. One "last updated" stamp for the table would be a false claim.
Freshness is per experiment.
**Cells in the same grid are not the same age, and nothing bounds the gap.** The
benchmark stopped running on a schedule on 1 October 2026: a run is dispatched against
a piece of work, so a column is as old as the last run that covered it. Until then the
frontier agents and the weak pair ran weekly and the `-no-skills` twins monthly, which
capped the gap at four weeks. There is no cap now. One "last updated" stamp for the
table would be a false claim; freshness is per experiment.

## The experiments, and what the pairing actually showed

Expand Down
2 changes: 1 addition & 1 deletion apps/framework/harness/run-eval.ts
Original file line number Diff line number Diff line change
Expand Up @@ -834,7 +834,7 @@ async function main() {
eval: ev.id,
// When this run finished. Recorded per run rather than per
// refresh, because the published grid is not one snapshot: the
// `-no-skills` twins refresh monthly and everything else weekly,
// a dispatched run covers only the experiments it asked for,
// and a targeted re-run replaces some pairs and leaves the rest,
// so two cells side by side can be a month apart. Without this
// the page can only describe the schedule, which is not the same
Expand Down
4 changes: 2 additions & 2 deletions apps/framework/lib/provenance.ts
Original file line number Diff line number Diff line change
Expand Up @@ -12,8 +12,8 @@ import {
* from a current one.
*
* `export-results --merge` carries forward any cell a run did not re-execute.
* That earns its place — the `-no-skills` twins refresh monthly and everything
* else weekly, so without it a weekly snapshot would have holes. What it lacks
* That earns its place — a dispatched run covers the experiments that run asked
* for, so without it a partial run's snapshot would have holes. What it lacks
* is any notion of a row going *out of date*. On 25 August the published file
* held 114 rows across six execution dates spanning thirteen days, 22 of them
* from 13 August: measured before the sandbox CLI moved to 2.5.0, before fixed
Expand Down
4 changes: 2 additions & 2 deletions apps/framework/scripts/export-results.ts
Original file line number Diff line number Diff line change
Expand Up @@ -75,8 +75,8 @@ const MERGE = rawArgs.includes('--merge');
*
* Opt-in, because a snapshot with holes and a snapshot with silent thirteen-day-
* old rows are both wrong and which is less wrong depends on what the snapshot
* is for. A release cut to report a measured change wants this on; a weekly
* refresh keeping the page populated probably does not. Without it the report
* is for. A release cut to report a measured change wants this on; a partial
* run keeping the page populated probably does not. Without it the report
* below still prints, so the staleness is visible either way — which is the
* actual defect in #60. Nothing was ever *said*.
*/
Expand Down
10 changes: 6 additions & 4 deletions reference/design-tokens.md
Original file line number Diff line number Diff line change
Expand Up @@ -127,7 +127,9 @@ because a flat delta honestly reported is worth more than an implied one, but do
build the hierarchy around it. The spread that carries signal is model capability: the
frontier agents pass nearly everything and a deliberately weaker model fails several.

**Cells in one grid are not the same age.** Frontier agents and the weak pair run
weekly; the `-no-skills` twins run monthly. A single "last updated" stamp for the table
would be wrong by up to four weeks on half the columns. Freshness is per experiment.
The design has no treatment for this because the cadence was set after it was drawn.
**Cells in one grid are not the same age, and since 1 October 2026 there is no bound
on how far apart they can be.** The benchmark came off the schedule, so cells are
measured when a run is dispatched against a piece of work. A single "last updated"
stamp for the table would be a false claim; freshness is per experiment. This used to
be "up to four weeks" — a weekly against a monthly — and that ceiling is gone, which
makes the gap worse rather than resolved. The design has no treatment for it.
Loading