Skip to content

Stop running the benchmark on a schedule; keep a regression run that tells us - #86

Merged
leggetter merged 4 commits into
mainfrom
regression-only-schedule
Oct 1, 2026
Merged

leggetter merged 4 commits into
mainfrom
regression-only-schedule

Conversation

@leggetter

@leggetter leggetter commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

The monthly was due to fire at 08:00 UTC today. The workflow is disabled right now so it did not; this PR is what replaces it.

What changes

Before After
Weekly Benchmark, 4 experiments, 76 cells, ~$185/month Regression suite, 6 experiments, 3 scenarios, 18 cells — cost and duration not yet measured
Monthly Full matrix, 114 cells, $90-110 Gone. Dispatched against a bucket of work instead
On failure A red run nobody is looking at Opens or comments on an issue labelled regression-alert

Do not enable on merge

Review found that the regression suite's green state has never been measured. The only record is one row in regression-eval-results.json: a failing cell from 10 August against claude-code-sonnet-5-docs-only, an experiment that no longer exists. "Every agent passing is the expected state" is an assumption, and the one measurement on record contradicts it — three of the six arms carry no skills, and all three scenarios are negative checks, which is the shape a skill-less arm fails.

Since this PR makes a failing cell fail the job, enabling blind risks an alert on week one, which teaches people to ignore the channel. Dispatch one regression run first (suite=regression, experiment_suite=regression, runs=2), record the pass state, cost and wall clock, then enable.

Why

The question "what decision does this run inform?" had stopped having an answer. Eight product findings are open and none has shipped, so the runs were widening a queue nobody was consuming. Five weeklies produced one harness defect (#79) and a lot of variance data about a delta we already know we cannot measure precisely enough (#2).

Two buckets, not a calendar

A benchmark run is dispatched against:

  • Eval changes — a scenario, a scorer, the base prompt, the CLI pin, a new model. The run measures what the change did to what we measure.
  • Product changes — a fix, docs change or skill change this benchmark found. The run measures whether it worked, which is the loop closing.

A run belonging to neither is buying data nobody has a decision waiting on.

Details worth reviewing

  • Two attempts per cell on the regression run. A guarded mistake failing once is either the mistake returning or noise; the second attempt is cheaper than a person adjudicating.
  • One alert issue, not one per week. A suite that breaks and stays broken should read as one problem getting worse.
  • The alert body leads with "the harness, not a regression", because that has been the cause every time so far, and tells the reader to open the transcript first.
  • Regression results do not reach the page — it renders the benchmark suite only — so nothing a reader sees moves.
  • The workflow stays disabled until this merges; note that disabling blocks workflow_dispatch too, so AGENTS.md records how to re-enable before a dispatched run.

🤖 Generated with Claude Code

https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK

leggetter and others added 2 commits October 1, 2026 08:32
…tells us

The benchmark ran weekly and monthly from August, about $185 a month plus
$90-110 a matrix. Paused because the question "what decision does this run
inform?" had stopped having an answer: eight product findings were open and
none had shipped, so the runs were widening a queue nobody was consuming. Five
weeklies produced one harness defect and a great deal of variance data about a
delta we already know we cannot measure precisely enough.

What stays automated is the regression suite — three scenarios guarding
mistakes already seen and fixed, every experiment, no environment to stand up.
Under $5 and under ten minutes. Two attempts per cell, because a regression
scenario that fails once is either the mistake returning or noise, and paying
for the second attempt is cheaper than a person deciding which.

And it now says something when it fails. A failing scheduled run was visible
only to somebody who went looking: the 21 September failure sat unnoticed for
four days until it was asked about. On failure the run opens an issue labelled
`regression-alert`, or comments on the open one, so a suite that breaks and
stays broken reads as one problem getting worse rather than four identical
issues nobody closes. The body says what a regression failure usually is — the
harness, not a regression — and to read the transcript first.

Benchmark runs are now dispatched against a bucket of work: an eval change, or
a product change this benchmark found. A run belonging to neither is buying
data nobody has a decision waiting on. AGENTS.md carries that, and the note
that the workflow is disabled so dispatching needs enabling first.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK
Review found the cadence change could never have worked. Correcting it, plus
the stale prose it left behind.

The scheduled branch asked for `suite=regression` with
`experiment_suite=benchmark,no-skills`, and a regression eval maps only to the
`regression` experiment suite — so every pair was skipped, `prepare` exited 1,
and the alert would have fired every Monday on a run that measured nothing.
Verified the fix produces 18 pairs.

The weak pair was named in a hardcoded six-experiment override while carrying
no `regression` suite, so it looked included and was not. They now carry it:
that pair holds nearly every failure in the suite, which is exactly where a
guarded mistake returning would show up first. The override is gone — the
suite decides who runs.

`regression-alert` had no checkout and no `GH_REPO`, so `gh` had no repository
to resolve against and the notifier would have failed silently. Its body was
indented inside the `run:` block, which GitHub renders as a code block; it is
now a heredoc de-indented to column zero, checked by running it.

A failing scenario did not fail the job, so the alert could not fire for the
thing it exists to detect. On the regression suite every agent passing is the
expected state, so a failing check now fails the job — scheduled runs only, and
the benchmark's semantics are untouched, where a failure is a score.

A regression-only run would also have republished the benchmark snapshot:
`publish-snapshot` writes a fresh `latest.json` from an untouched file, so the
`runId` would have named a run that produced none of its rows, against a
contract README states explicitly. Gated on benchmark pairs being present.

And the claims. "Under $5 and under ten minutes" was a planning estimate from
before any regression run existed, promoted to fact in AGENTS.md; the measured
3.4 minutes per cell puts eighteen cells nearer an hour, so both places now say
it is unmeasured. Eleven stale cadence statements across AGENTS.md, LOOPS.md,
the delivery plan and this workflow's own comments are corrected, including the
one in Status that an agent reads before anything else.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK
Comment thread .github/workflows/eval-refresh.yml Outdated
exit 1
fi

- name: Fail on a regression that came back

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

poor name

leggetter and others added 2 commits October 1, 2026 11:46
"Fail on a regression that came back" names one of the two things a failure
here means, and the less likely one — the step fails for a broken harness far
more often than for a guarded mistake returning, which is why the alert body
leads with that. It also reads as a wish rather than a check.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK
Four from the second review, all small, one of them a correction to a claim I
made in a commit message.

The new check keyed off `github.event_name == 'schedule'` while its own comment
said the distinction was benchmark-versus-regression. Those are the same thing
only while the single cron happens to be regression-only; add a benchmark cron
back and every benchmark cell that merely scored a failure would fail its job.
It now keys on `matrix.eval_suite == 'regression'`, which is what the comment
meant, with the schedule clause kept because a dispatched run has somebody
reading it.

The alert body's `sed` stripped twelve spaces. YAML's block scalar had already
removed ten, so the lines arrived with two and the `sed` matched nothing. It
rendered correctly anyway, because two is below the four that makes a markdown
code block — so the mechanism was dead and the output was fine by luck. The
earlier commit message said this was "de-indented to column zero, checked by
running it": I did run it, in a standalone script where the indent was twelve,
which is not what YAML hands to bash. Now `sed 's/^  //'`, checked by parsing
the workflow and running the body exactly as the runner would.

The alert could not fire for a `publish-results` failure, which is a third of
the pipeline and has failed on its own twice — a shallow clone that could not
rebase, and a re-run against a moved branch (#74). Nor for a cancelled run,
which is what the provider-dead path produces when credits run out. Both are
covered now.

And the "$5 and under 10 minutes" estimate this branch says it retired was
still in the delivery plan, two lines below the bullet that was edited.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK
@leggetter
leggetter merged commit 0fb4012 into main Oct 1, 2026
1 check passed
@leggetter
leggetter deleted the regression-only-schedule branch October 1, 2026 11:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant