Skip to content
Draft
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
95 changes: 58 additions & 37 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,52 +32,73 @@ for e in sorted(m): print(e, sum(1 for v in m[e].values() if v), '/', len(m[e]))
`['hookdeck', 'event-gateway']`, `-no-skills` twins of each, and
`codex-gpt-5.4-mini` in both arms as a deliberately weaker model.

**What the numbers say.** Nine of nineteen scenarios discriminate, which is a
healthy benchmark rather than a flat one. The frontier agents pass nearly
everything; the weak model is where most failures live, which is the floor
working as intended.

The skills axis is the interesting result and it is not uniform. Claude gains
one scenario from skills, GPT-5.6 nets zero, and the weak model is **two worse
with skills than without** — on the most recent published run it loses four
scenarios and gains two. That direction is a finding about our documentation
rather than about the model, and there is a known mechanism: a skill that lists
example values is read as an exhaustive list, which once led a weak model to
conclude a supported provider was unsupported. Do not report the skills delta as
a single number; it has a different sign at different capability levels.

**The weak-model figure is −2 and was written here as −3 for eleven days.** Both
−3 readings are from 13 August and no run since has reproduced them: eight of
the nine later runs read −2 and one read 0. 13 August is also the credit-outage
day described under costs below, so treat anything measured that day as suspect
until its rows are checked against the outage window — the point of that
paragraph is that a bad population is identified by its cause, and this figure
was never re-derived after the cause was found. Recompute from `results/runs/`
rather than quoting this paragraph:
**What the numbers say.** Seven of the nineteen scenarios are failed by at least
one experiment on the published snapshot, which is a healthy benchmark rather
than a flat one. The frontier agents pass nearly everything; the weak model is
where most failures live, which is the floor working as intended.

**Read a skills delta only off two arms that ran the same day.**
`results/latest.json` is a merged snapshot (#60): the frontier `-no-skills` arms
in it were executed on 1 September and their `+skills` twins on 14 September, so
subtracting them compares two instruments a fortnight apart and does not produce
a delta at all. Only the weak pair runs weekly, so it is the only clean
comparison in that file, and there it reads **+4** — 18/19 with skills against
14/19 without.

**The weak model is not worse with skills on any evidence since August, and this
paragraph said it was for six weeks.** Across the five full-suite pairwise
executions of that pair the delta reads −3, −2, −1, +1, +4, with nothing in the
instrument changing between the last three. #2 was closed on 22 September on that
basis. Pooling the three September executions gives `+skills` 48/57 against
44/57, and all of the gain is Outpost — 14/15 against 10/15, with Event Gateway
level at 34/42 in both arms. A quantity that has taken five values on an
unchanged instrument is not a measurement yet, so do not quote a weak-model
skills delta without saying which executions produced it.

The −3 is from 13 August, the credit-outage day described under costs below, so
treat anything measured that day as suspect until its rows are checked against
the outage window: a bad population is identified by its cause, not by what its
rows look like.

**Recompute per execution, not per snapshot file.** A snapshot republishes the
cells a run did not re-measure, so iterating `results/runs/*.json` counts one
execution many times — the 17 August pair appears in six files and reads −2 in
all six, which is how "eight of the nine later runs read −2" came to be written
here. Key each cell by its run before counting anything:

```bash
python3 -c "
import json,glob,os
import json,glob
from collections import defaultdict
cells={}
for f in sorted(glob.glob('results/runs/*.json')):
arms=defaultdict(lambda:[0,0])
for r in json.load(open(f))['results']:
e=r['experiment']
if not e.startswith('codex-gpt-5.4-mini'): continue
arms['base' if e.endswith('-no-skills') else 'skills'][0] += 1 if r['passed'] else 0
arms['base' if e.endswith('-no-skills') else 'skills'][1] += 1
s,b=arms['skills'],arms['base']
if s[1] and b[1]: print(os.path.basename(f)[:16], f'{s[0]}/{s[1]}', f'{b[0]}/{b[1]}', f'{s[0]-b[0]:+d}')
run = r.get('runId') or (r.get('ranAt') or '')[:10]
cells[(run, r['eval'], r['experiment'])] = r
arms=defaultdict(lambda: defaultdict(lambda: [0, 0]))
for (run, _, exp), r in cells.items():
if not exp.startswith('codex-gpt-5.4-mini'): continue
a = 'base' if exp.endswith('-no-skills') else 'skills'
arms[run][a][0] += 1 if r['passed'] else 0
arms[run][a][1] += 1
for run in sorted(arms, key=str):
s, b = arms[run]['skills'], arms[run]['base']
if s[1] and b[1]: print(f'{str(run):14} {s[0]}/{s[1]:<3} {b[0]}/{b[1]:<3} {s[0]-b[0]:+d}')
"
```

**Do not read the Outpost-only +2 as contradicting it.** Loop 2's corrected
measurement was `+skills` 12/12 against `-no-skills` 10/12 — but that is five
Outpost scenarios across all three models, not fifteen-plus scenarios on the
weak model, and the two populations answer different questions. They were
conflated once in conversation into a claim that the delta's *sign* had flipped,
which no run supports. State the scenario set and the model with any skills
delta, or it will be compared against a number measuring something else.
Two things that output will not tell you. The denominators are the filter: a row
reading `2/2` is a partial re-run of a couple of cells, not a suite delta. And
`runId` only exists from 1 September, so earlier executions fall back to their
date and two on the same day merge — which is another reason 13 August needs
reading by cause rather than by row.

**State the scenario set and the model with any skills delta**, or it will be
compared against a number measuring something else. Loop 2's corrected
measurement was `+skills` 12/12 against `-no-skills` 10/12, but that is five
Outpost scenarios across all three models rather than nineteen scenarios on the
weak model, and the two populations answer different questions. That they now
point the same way is not a reason to merge them.

**Product findings come from runs, not speculation.** Several concern
`hookdeck listen`: it crashes without a TTY unless given `--output compact`; a
Expand Down
Loading