From c432bf576d6ba42cda7752c9cfea96779902d534 Mon Sep 17 00:00:00 2001 From: Phil Leggetter Date: Mon, 21 Sep 2026 14:25:10 +0100 Subject: [PATCH] Say what justifies a release, and what does not MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The convention said "cut one when there is a measured change to report" and left the three cases anyone actually faces unanswered: time passing, the benchmark changing, the product changing. Three triggers, any one sufficient. The product moved — something shipped or a finding came out of a run. The instrument changed what it measures and a full run has been measured under it, so the numbers and the notes describing them agree. Or the snapshot has gone stale, which has a deadline attached: transcripts expire at ninety days, so a snapshot nobody released is evidence nobody can check later. And one explicit non-trigger, because it is the tempting one. A run whose numbers moved is not a reason. The weak model's skills delta read -1 on 1 September and +4 on 14 September with nothing changed between them; eleven cells flipped and two of them moved in both directions across the arms of the same pair. Releasing on movement publishes variance and spends the changelog on non-events. Also records that the monthly matrix is the better thing to package: a weekly covers four experiments, so its snapshot carries the frontier -no-skills arms forward and a release calling it "one run" would be wrong. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01MQzUoMAwEBJWpEGVvVzSjK --- AGENTS.md | 26 ++++++++++++++++++++++++++ 1 file changed, 26 insertions(+) diff --git a/AGENTS.md b/AGENTS.md index cd17099..1877702 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -221,6 +221,32 @@ cannot separate a fix from variance. A release is how an improvement is published. It is not a calendar event: cut one when there is a measured change to report, not on a schedule. +**What justifies one.** The page is release-pinned, so cutting a release is the act of +changing what the public sees — and not cutting one is how a run gets read before it is +believed. Any of these three is enough: + +1. **The product moved.** Something shipped that this benchmark found, or a finding + worth publishing came out of a run. This is the loop working and it is what the + changelog is for. +2. **The instrument changed what it measures**, and a full run has been measured under + it. A scenario, a scorer, the base prompt, the sandbox CLI pin. Release so the + published numbers and the notes describing them agree; until then the page is + showing numbers measured under something the notes do not describe. +3. **The published snapshot has gone stale** — roughly six weeks, or sooner if the + instrument has moved under it. Transcripts expire at ninety days (#21), so a + snapshot nobody released is evidence nobody can check later. + +**What does not justify one: a run whose numbers moved.** Movement is the default. The +weak model's skills delta read -1 on 1 September and +4 on 14 September with no change +to the instrument between them, and eleven cells flipped, two of them in both +directions across the arms of the same pair. Releasing on movement publishes variance +and spends the changelog on non-events. + +**Prefer packaging the monthly matrix.** A weekly covers four experiments, so its +snapshot carries the frontier `-no-skills` arms forward from whenever they last ran — +publishable, but it means a release whose notes say "one run" would be wrong. The +monthly measures all six in one pass. + **A release represents one run and what changed since the last one.** Tag the results commit, so the release points at exactly the data it describes, and `results/runs/.json` is the immutable snapshot behind it.