Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 7 additions & 5 deletions .github/workflows/eval-refresh.yml
Original file line number Diff line number Diff line change
Expand Up @@ -17,11 +17,13 @@ name: Refresh eval results
# instead of going red where nobody looks. The 21 September failure sat
# unnoticed until somebody asked about it four days later.
#
# Cost and duration are not yet measured. The $5-and-ten-minutes figure that
# used to sit here came from `.plans/delivery-plan.md` at planning time, before
# any regression run existed; measured wall clock on the benchmark is about
# 3.4 minutes per cell serialised, which would put eighteen cells nearer an
# hour. Replace this sentence with a measurement after the first run.
# Measured on the first run, 1 October 2026 (run 36869979661): eighteen cells
# green at two attempts each, fifty minutes of wall clock. That is whole-run
# time including runner queueing, not per-cell — the cells run in parallel, so
# no per-cell figure can be read off it. The $5-and-ten-minutes estimate that
# used to sit here was a planning figure from before any regression run
# existed; cost is still unmeasured, because nothing in the harness records it
# for an agent that reports tokens rather than dollars.
#
# **Benchmark runs are dispatched against a bucket of work, not a calendar.**
# Something changed what we measure, or something shipped to the product, and
Expand Down
51 changes: 31 additions & 20 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -37,14 +37,19 @@ healthy benchmark rather than a flat one. The frontier agents pass nearly
everything; the weak model is where most failures live, which is the floor
working as intended.

The skills axis is the interesting result and it is not uniform. Claude gains
one scenario from skills, GPT-5.6 nets zero, and the weak model is **two worse
with skills than without** — on the most recent published run it loses four
scenarios and gains two. That direction is a finding about our documentation
rather than about the model, and there is a known mechanism: a skill that lists
example values is read as an exhaustive list, which once led a weak model to
conclude a supported provider was unsupported. Do not report the skills delta as
a single number; it has a different sign at different capability levels.
The skills axis has not settled, and this paragraph has been rewritten three times
in the direction of less confidence. It once said the weak model was **two worse**
with skills than without; on 28 September that model reads **four better** (18/19
against 14/19), and across the four September runs the delta has read -1, +1, +4
and +4. #2 was closed on 22 September with its premise withdrawn: the -3 that
started it is from the credit-outage day and no run since reproduced it.

What survives is the warning rather than the number. **Do not report the skills
delta as a single figure**, do not report a sign as replicated until it has
replicated, and read it per row — the disagreements move between runs, so a total
hides which cells produced it. The one mechanism that is documented rather than
inferred stands: a skill that lists example values is read as an exhaustive list,
which once led a weak model to conclude a supported provider was unsupported.

**The weak-model figure is −2 and was written here as −3 for eleven days.** Both
−3 readings are from 13 August and no run since has reproduced them: eight of
Expand Down Expand Up @@ -98,8 +103,7 @@ others would have caught.

`eval-refresh` runs the **regression suite** weekly (Monday 06:00 UTC, every
experiment) and nothing else on a schedule. Benchmark runs are dispatched
against a bucket of work — see Runs below. The workflow is disabled, so a
dispatch needs `gh workflow enable eval-refresh.yml` first. Check `gh secret list` against the workflow env
against a bucket of work — see Runs below. Check `gh secret list` against the workflow env
rather than trusting any list written here: `OUTPOST_API_KEY` was documented as
a secret before it existed, and the first full matrix run scored `outpost-001`
as six agent failures because of it.
Expand Down Expand Up @@ -173,11 +177,15 @@ state, so a failing check **fails the job** here — the opposite of the benchma
where a failure is a score — and the run opens or updates an issue labelled
`regression-alert`, because a red run in a tab nobody has open is not a notification.

Its cost and duration are **not yet measured**. The "$5 and ten minutes" figure that
circulated came from the delivery plan at planning time, before any regression run
existed, and the benchmark's measured 3.4 minutes per cell serialised would put
eighteen cells nearer an hour. Measure it after the first run rather than quoting the
estimate again.
**Measured on the first run**, 1 October 2026 ([run 36869979661](https://github.com/hookdeck/evals/actions/runs/36869979661)): eighteen cells
green at two attempts each, **fifty minutes** of wall clock. Whole-run time including
runner queueing — the cells run in parallel, so there is no per-cell figure in it. The
"$5 and ten minutes" that circulated was a planning estimate from before any regression
run existed. Cost is still unmeasured.

That run is also the first evidence that every agent passing is the expected state here.
It was an assumption until then, and the only record that existed contradicted it: one
failing row from 10 August against an experiment that no longer exists.

**A benchmark run is dispatched against a bucket of work, never a date.** Two buckets,
and a run belongs to one of them:
Expand All @@ -194,11 +202,13 @@ consuming, at about $185 a month plus $90-110 a matrix. Five weeklies had produc
harness defect and a lot of variance data about a delta we already know we cannot
measure precisely enough (#2).

Re-enable the workflow before dispatching: it is disabled, which blocks
`workflow_dispatch` as well as the cron.
The workflow was disabled between 1 October 07:30 and 13:30 UTC so the monthly matrix
could not fire while the schedule was being changed. It is active again. Disabling is
the way to stop a cron without a commit, and it blocks `workflow_dispatch` too:

```bash
gh workflow enable eval-refresh.yml
gh workflow disable eval-refresh.yml # stops the cron and dispatch
gh workflow enable eval-refresh.yml # both again
```

## What to work on next
Expand Down Expand Up @@ -252,8 +262,9 @@ Two rules that are the whole point of the file:
- **Record the loops that failed.** A change that did not work is more informative
than one that did, and omitting them makes the rest less believable.

Re-runs for a loop need at least three attempts. A dispatched run defaults to one and
cannot separate a fix from variance.
Re-runs for a loop need at least three attempts. A dispatched run defaults to two,
which is enough to stop one unlucky cell deciding a verdict and not enough to separate
a fix from variance — ask for three.

## Releases

Expand Down
Loading