Skip to content

Record what the first regression run measured, and fix what it made stale - #91

Open
leggetter wants to merge 1 commit into
mainfrom
record-regression-measurement
Open

leggetter wants to merge 1 commit into
mainfrom
record-regression-measurement

Conversation

@leggetter

Copy link
Copy Markdown
Collaborator

The run #86 asked for before the schedule stays enabled.

Measured

Run 36869979661, 1 October: eighteen cells green at two attempts each, fifty minutes wall clock, regression-alert correctly skipped.

Two things that figure is not: it is whole-run time including runner queueing, and the cells run in parallel, so no per-cell number can be read off it. Cost remains unmeasured — nothing records dollars for an agent that reports tokens.

It is also the first evidence for the assumption the whole design rests on. "Every agent passing is the expected state" was unverified until today, and the only record that existed contradicted it: one failing row from 10 August against an experiment that no longer exists.

Three claims the change made wrong

Said Is
"The workflow is disabled, so a dispatch needs gh workflow enable first" Active. It was off for six hours this morning so the monthly could not fire mid-change
"A dispatched run defaults to one" Two, set in #86
Status: the weak model is "two worse with skills than without" On 28 September it is four better (18/19 against 14/19)

The Status paragraph, rewritten a third time

It has now been revised three times, each towards less confidence, so it carries the warning instead of a number: do not report the skills delta as a single figure, do not call a sign replicated until it has, and read it per row because the disagreements move between runs. #2 was closed on 22 September with its premise withdrawn.

🤖 Generated with Claude Code

https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK

…tale

The run that #86 said to take before leaving the schedule enabled: eighteen
cells green at two attempts, fifty minutes of wall clock, alert job correctly
skipped. Whole-run time including queueing — the cells run in parallel, so
there is no per-cell figure in it, and cost is still unmeasured because nothing
records dollars for an agent that reports tokens.

That run is also the first evidence for the assumption the design rests on.
"Every agent passing is the expected state" was unverified until today, and the
only record that existed contradicted it: one failing row from 10 August
against an experiment that no longer exists.

Three statements the change made wrong. The workflow is active, not disabled —
it was off for six hours this morning so the monthly could not fire mid-change,
and saying "it is disabled" would send the next reader to enable something
already enabled. A dispatched run defaults to two attempts now, not one. And
the Status section said the weak model is two worse with skills than without;
on 28 September it is four better, and across four September runs the delta
reads -1, +1, +4, +4.

That last paragraph has now been rewritten three times, each time towards less
confidence, so it now carries the warning rather than a number: do not report
the delta as a single figure, do not call a sign replicated until it has
replicated, and read it per row because the disagreements move.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant