From 3d8c16fd30f85a7d0d1257c155c8821f80879861 Mon Sep 17 00:00:00 2001 From: Phil Leggetter Date: Thu, 1 Oct 2026 15:35:06 +0100 Subject: [PATCH] Record what the first regression run measured, and fix what it made stale MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The run that #86 said to take before leaving the schedule enabled: eighteen cells green at two attempts, fifty minutes of wall clock, alert job correctly skipped. Whole-run time including queueing — the cells run in parallel, so there is no per-cell figure in it, and cost is still unmeasured because nothing records dollars for an agent that reports tokens. That run is also the first evidence for the assumption the design rests on. "Every agent passing is the expected state" was unverified until today, and the only record that existed contradicted it: one failing row from 10 August against an experiment that no longer exists. Three statements the change made wrong. The workflow is active, not disabled — it was off for six hours this morning so the monthly could not fire mid-change, and saying "it is disabled" would send the next reader to enable something already enabled. A dispatched run defaults to two attempts now, not one. And the Status section said the weak model is two worse with skills than without; on 28 September it is four better, and across four September runs the delta reads -1, +1, +4, +4. That last paragraph has now been rewritten three times, each time towards less confidence, so it now carries the warning rather than a number: do not report the delta as a single figure, do not call a sign replicated until it has replicated, and read it per row because the disagreements move. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK --- .github/workflows/eval-refresh.yml | 12 ++++--- AGENTS.md | 51 ++++++++++++++++++------------ 2 files changed, 38 insertions(+), 25 deletions(-) diff --git a/.github/workflows/eval-refresh.yml b/.github/workflows/eval-refresh.yml index b40d931..d5dd73b 100644 --- a/.github/workflows/eval-refresh.yml +++ b/.github/workflows/eval-refresh.yml @@ -17,11 +17,13 @@ name: Refresh eval results # instead of going red where nobody looks. The 21 September failure sat # unnoticed until somebody asked about it four days later. # -# Cost and duration are not yet measured. The $5-and-ten-minutes figure that -# used to sit here came from `.plans/delivery-plan.md` at planning time, before -# any regression run existed; measured wall clock on the benchmark is about -# 3.4 minutes per cell serialised, which would put eighteen cells nearer an -# hour. Replace this sentence with a measurement after the first run. +# Measured on the first run, 1 October 2026 (run 36869979661): eighteen cells +# green at two attempts each, fifty minutes of wall clock. That is whole-run +# time including runner queueing, not per-cell — the cells run in parallel, so +# no per-cell figure can be read off it. The $5-and-ten-minutes estimate that +# used to sit here was a planning figure from before any regression run +# existed; cost is still unmeasured, because nothing in the harness records it +# for an agent that reports tokens rather than dollars. # # **Benchmark runs are dispatched against a bucket of work, not a calendar.** # Something changed what we measure, or something shipped to the product, and diff --git a/AGENTS.md b/AGENTS.md index 51380bc..107b7db 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -37,14 +37,19 @@ healthy benchmark rather than a flat one. The frontier agents pass nearly everything; the weak model is where most failures live, which is the floor working as intended. -The skills axis is the interesting result and it is not uniform. Claude gains -one scenario from skills, GPT-5.6 nets zero, and the weak model is **two worse -with skills than without** — on the most recent published run it loses four -scenarios and gains two. That direction is a finding about our documentation -rather than about the model, and there is a known mechanism: a skill that lists -example values is read as an exhaustive list, which once led a weak model to -conclude a supported provider was unsupported. Do not report the skills delta as -a single number; it has a different sign at different capability levels. +The skills axis has not settled, and this paragraph has been rewritten three times +in the direction of less confidence. It once said the weak model was **two worse** +with skills than without; on 28 September that model reads **four better** (18/19 +against 14/19), and across the four September runs the delta has read -1, +1, +4 +and +4. #2 was closed on 22 September with its premise withdrawn: the -3 that +started it is from the credit-outage day and no run since reproduced it. + +What survives is the warning rather than the number. **Do not report the skills +delta as a single figure**, do not report a sign as replicated until it has +replicated, and read it per row — the disagreements move between runs, so a total +hides which cells produced it. The one mechanism that is documented rather than +inferred stands: a skill that lists example values is read as an exhaustive list, +which once led a weak model to conclude a supported provider was unsupported. **The weak-model figure is −2 and was written here as −3 for eleven days.** Both −3 readings are from 13 August and no run since has reproduced them: eight of @@ -98,8 +103,7 @@ others would have caught. `eval-refresh` runs the **regression suite** weekly (Monday 06:00 UTC, every experiment) and nothing else on a schedule. Benchmark runs are dispatched -against a bucket of work — see Runs below. The workflow is disabled, so a -dispatch needs `gh workflow enable eval-refresh.yml` first. Check `gh secret list` against the workflow env +against a bucket of work — see Runs below. Check `gh secret list` against the workflow env rather than trusting any list written here: `OUTPOST_API_KEY` was documented as a secret before it existed, and the first full matrix run scored `outpost-001` as six agent failures because of it. @@ -173,11 +177,15 @@ state, so a failing check **fails the job** here — the opposite of the benchma where a failure is a score — and the run opens or updates an issue labelled `regression-alert`, because a red run in a tab nobody has open is not a notification. -Its cost and duration are **not yet measured**. The "$5 and ten minutes" figure that -circulated came from the delivery plan at planning time, before any regression run -existed, and the benchmark's measured 3.4 minutes per cell serialised would put -eighteen cells nearer an hour. Measure it after the first run rather than quoting the -estimate again. +**Measured on the first run**, 1 October 2026 ([run 36869979661](https://github.com/hookdeck/evals/actions/runs/36869979661)): eighteen cells +green at two attempts each, **fifty minutes** of wall clock. Whole-run time including +runner queueing — the cells run in parallel, so there is no per-cell figure in it. The +"$5 and ten minutes" that circulated was a planning estimate from before any regression +run existed. Cost is still unmeasured. + +That run is also the first evidence that every agent passing is the expected state here. +It was an assumption until then, and the only record that existed contradicted it: one +failing row from 10 August against an experiment that no longer exists. **A benchmark run is dispatched against a bucket of work, never a date.** Two buckets, and a run belongs to one of them: @@ -194,11 +202,13 @@ consuming, at about $185 a month plus $90-110 a matrix. Five weeklies had produc harness defect and a lot of variance data about a delta we already know we cannot measure precisely enough (#2). -Re-enable the workflow before dispatching: it is disabled, which blocks -`workflow_dispatch` as well as the cron. +The workflow was disabled between 1 October 07:30 and 13:30 UTC so the monthly matrix +could not fire while the schedule was being changed. It is active again. Disabling is +the way to stop a cron without a commit, and it blocks `workflow_dispatch` too: ```bash -gh workflow enable eval-refresh.yml +gh workflow disable eval-refresh.yml # stops the cron and dispatch +gh workflow enable eval-refresh.yml # both again ``` ## What to work on next @@ -252,8 +262,9 @@ Two rules that are the whole point of the file: - **Record the loops that failed.** A change that did not work is more informative than one that did, and omitting them makes the rest less believable. -Re-runs for a loop need at least three attempts. A dispatched run defaults to one and -cannot separate a fix from variance. +Re-runs for a loop need at least three attempts. A dispatched run defaults to two, +which is enough to stop one unlucky cell deciding a verdict and not enough to separate +a fix from variance — ask for three. ## Releases