Stop running the benchmark on a schedule; keep a regression run that tells us - #86
Merged
Merged
Conversation
…tells us The benchmark ran weekly and monthly from August, about $185 a month plus $90-110 a matrix. Paused because the question "what decision does this run inform?" had stopped having an answer: eight product findings were open and none had shipped, so the runs were widening a queue nobody was consuming. Five weeklies produced one harness defect and a great deal of variance data about a delta we already know we cannot measure precisely enough. What stays automated is the regression suite — three scenarios guarding mistakes already seen and fixed, every experiment, no environment to stand up. Under $5 and under ten minutes. Two attempts per cell, because a regression scenario that fails once is either the mistake returning or noise, and paying for the second attempt is cheaper than a person deciding which. And it now says something when it fails. A failing scheduled run was visible only to somebody who went looking: the 21 September failure sat unnoticed for four days until it was asked about. On failure the run opens an issue labelled `regression-alert`, or comments on the open one, so a suite that breaks and stays broken reads as one problem getting worse rather than four identical issues nobody closes. The body says what a regression failure usually is — the harness, not a regression — and to read the transcript first. Benchmark runs are now dispatched against a bucket of work: an eval change, or a product change this benchmark found. A run belonging to neither is buying data nobody has a decision waiting on. AGENTS.md carries that, and the note that the workflow is disabled so dispatching needs enabling first. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK
Review found the cadence change could never have worked. Correcting it, plus the stale prose it left behind. The scheduled branch asked for `suite=regression` with `experiment_suite=benchmark,no-skills`, and a regression eval maps only to the `regression` experiment suite — so every pair was skipped, `prepare` exited 1, and the alert would have fired every Monday on a run that measured nothing. Verified the fix produces 18 pairs. The weak pair was named in a hardcoded six-experiment override while carrying no `regression` suite, so it looked included and was not. They now carry it: that pair holds nearly every failure in the suite, which is exactly where a guarded mistake returning would show up first. The override is gone — the suite decides who runs. `regression-alert` had no checkout and no `GH_REPO`, so `gh` had no repository to resolve against and the notifier would have failed silently. Its body was indented inside the `run:` block, which GitHub renders as a code block; it is now a heredoc de-indented to column zero, checked by running it. A failing scenario did not fail the job, so the alert could not fire for the thing it exists to detect. On the regression suite every agent passing is the expected state, so a failing check now fails the job — scheduled runs only, and the benchmark's semantics are untouched, where a failure is a score. A regression-only run would also have republished the benchmark snapshot: `publish-snapshot` writes a fresh `latest.json` from an untouched file, so the `runId` would have named a run that produced none of its rows, against a contract README states explicitly. Gated on benchmark pairs being present. And the claims. "Under $5 and under ten minutes" was a planning estimate from before any regression run existed, promoted to fact in AGENTS.md; the measured 3.4 minutes per cell puts eighteen cells nearer an hour, so both places now say it is unmeasured. Eleven stale cadence statements across AGENTS.md, LOOPS.md, the delivery plan and this workflow's own comments are corrected, including the one in Status that an agent reads before anything else. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK
leggetter
commented
Oct 1, 2026
| exit 1 | ||
| fi | ||
|
|
||
| - name: Fail on a regression that came back |
"Fail on a regression that came back" names one of the two things a failure here means, and the less likely one — the step fails for a broken harness far more often than for a guarded mistake returning, which is why the alert body leads with that. It also reads as a wish rather than a check. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK
Four from the second review, all small, one of them a correction to a claim I made in a commit message. The new check keyed off `github.event_name == 'schedule'` while its own comment said the distinction was benchmark-versus-regression. Those are the same thing only while the single cron happens to be regression-only; add a benchmark cron back and every benchmark cell that merely scored a failure would fail its job. It now keys on `matrix.eval_suite == 'regression'`, which is what the comment meant, with the schedule clause kept because a dispatched run has somebody reading it. The alert body's `sed` stripped twelve spaces. YAML's block scalar had already removed ten, so the lines arrived with two and the `sed` matched nothing. It rendered correctly anyway, because two is below the four that makes a markdown code block — so the mechanism was dead and the output was fine by luck. The earlier commit message said this was "de-indented to column zero, checked by running it": I did run it, in a standalone script where the indent was twelve, which is not what YAML hands to bash. Now `sed 's/^ //'`, checked by parsing the workflow and running the body exactly as the runner would. The alert could not fire for a `publish-results` failure, which is a third of the pipeline and has failed on its own twice — a shallow clone that could not rebase, and a re-run against a moved branch (#74). Nor for a cancelled run, which is what the provider-dead path produces when credits run out. Both are covered now. And the "$5 and under 10 minutes" estimate this branch says it retired was still in the delivery plan, two lines below the bullet that was edited. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK
This was referenced Oct 1, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The monthly was due to fire at 08:00 UTC today. The workflow is disabled right now so it did not; this PR is what replaces it.
What changes
regression-alertDo not enable on merge
Review found that the regression suite's green state has never been measured. The only record is one row in
regression-eval-results.json: a failing cell from 10 August againstclaude-code-sonnet-5-docs-only, an experiment that no longer exists. "Every agent passing is the expected state" is an assumption, and the one measurement on record contradicts it — three of the six arms carry no skills, and all three scenarios are negative checks, which is the shape a skill-less arm fails.Since this PR makes a failing cell fail the job, enabling blind risks an alert on week one, which teaches people to ignore the channel. Dispatch one regression run first (
suite=regression, experiment_suite=regression, runs=2), record the pass state, cost and wall clock, then enable.Why
The question "what decision does this run inform?" had stopped having an answer. Eight product findings are open and none has shipped, so the runs were widening a queue nobody was consuming. Five weeklies produced one harness defect (#79) and a lot of variance data about a delta we already know we cannot measure precisely enough (#2).
Two buckets, not a calendar
A benchmark run is dispatched against:
A run belonging to neither is buying data nobody has a decision waiting on.
Details worth reviewing
workflow_dispatchtoo, so AGENTS.md records how to re-enable before a dispatched run.🤖 Generated with Claude Code
https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK