diff --git a/AGENTS.md b/AGENTS.md index 1877702..879c48a 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -32,52 +32,73 @@ for e in sorted(m): print(e, sum(1 for v in m[e].values() if v), '/', len(m[e])) `['hookdeck', 'event-gateway']`, `-no-skills` twins of each, and `codex-gpt-5.4-mini` in both arms as a deliberately weaker model. -**What the numbers say.** Nine of nineteen scenarios discriminate, which is a -healthy benchmark rather than a flat one. The frontier agents pass nearly -everything; the weak model is where most failures live, which is the floor -working as intended. - -The skills axis is the interesting result and it is not uniform. Claude gains -one scenario from skills, GPT-5.6 nets zero, and the weak model is **two worse -with skills than without** — on the most recent published run it loses four -scenarios and gains two. That direction is a finding about our documentation -rather than about the model, and there is a known mechanism: a skill that lists -example values is read as an exhaustive list, which once led a weak model to -conclude a supported provider was unsupported. Do not report the skills delta as -a single number; it has a different sign at different capability levels. - -**The weak-model figure is −2 and was written here as −3 for eleven days.** Both -−3 readings are from 13 August and no run since has reproduced them: eight of -the nine later runs read −2 and one read 0. 13 August is also the credit-outage -day described under costs below, so treat anything measured that day as suspect -until its rows are checked against the outage window — the point of that -paragraph is that a bad population is identified by its cause, and this figure -was never re-derived after the cause was found. Recompute from `results/runs/` -rather than quoting this paragraph: +**What the numbers say.** Seven of the nineteen scenarios are failed by at least +one experiment on the published snapshot, which is a healthy benchmark rather +than a flat one. The frontier agents pass nearly everything; the weak model is +where most failures live, which is the floor working as intended. + +**Read a skills delta only off two arms that ran the same day.** +`results/latest.json` is a merged snapshot (#60): the frontier `-no-skills` arms +in it were executed on 1 September and their `+skills` twins on 14 September, so +subtracting them compares two instruments a fortnight apart and does not produce +a delta at all. Only the weak pair runs weekly, so it is the only clean +comparison in that file, and there it reads **+4** — 18/19 with skills against +14/19 without. + +**The weak model is not worse with skills on any evidence since August, and this +paragraph said it was for six weeks.** Across the five full-suite pairwise +executions of that pair the delta reads −3, −2, −1, +1, +4, with nothing in the +instrument changing between the last three. #2 was closed on 22 September on that +basis. Pooling the three September executions gives `+skills` 48/57 against +44/57, and all of the gain is Outpost — 14/15 against 10/15, with Event Gateway +level at 34/42 in both arms. A quantity that has taken five values on an +unchanged instrument is not a measurement yet, so do not quote a weak-model +skills delta without saying which executions produced it. + +The −3 is from 13 August, the credit-outage day described under costs below, so +treat anything measured that day as suspect until its rows are checked against +the outage window: a bad population is identified by its cause, not by what its +rows look like. + +**Recompute per execution, not per snapshot file.** A snapshot republishes the +cells a run did not re-measure, so iterating `results/runs/*.json` counts one +execution many times — the 17 August pair appears in six files and reads −2 in +all six, which is how "eight of the nine later runs read −2" came to be written +here. Key each cell by its run before counting anything: ```bash python3 -c " -import json,glob,os +import json,glob from collections import defaultdict +cells={} for f in sorted(glob.glob('results/runs/*.json')): - arms=defaultdict(lambda:[0,0]) for r in json.load(open(f))['results']: - e=r['experiment'] - if not e.startswith('codex-gpt-5.4-mini'): continue - arms['base' if e.endswith('-no-skills') else 'skills'][0] += 1 if r['passed'] else 0 - arms['base' if e.endswith('-no-skills') else 'skills'][1] += 1 - s,b=arms['skills'],arms['base'] - if s[1] and b[1]: print(os.path.basename(f)[:16], f'{s[0]}/{s[1]}', f'{b[0]}/{b[1]}', f'{s[0]-b[0]:+d}') + run = r.get('runId') or (r.get('ranAt') or '')[:10] + cells[(run, r['eval'], r['experiment'])] = r +arms=defaultdict(lambda: defaultdict(lambda: [0, 0])) +for (run, _, exp), r in cells.items(): + if not exp.startswith('codex-gpt-5.4-mini'): continue + a = 'base' if exp.endswith('-no-skills') else 'skills' + arms[run][a][0] += 1 if r['passed'] else 0 + arms[run][a][1] += 1 +for run in sorted(arms, key=str): + s, b = arms[run]['skills'], arms[run]['base'] + if s[1] and b[1]: print(f'{str(run):14} {s[0]}/{s[1]:<3} {b[0]}/{b[1]:<3} {s[0]-b[0]:+d}') " ``` -**Do not read the Outpost-only +2 as contradicting it.** Loop 2's corrected -measurement was `+skills` 12/12 against `-no-skills` 10/12 — but that is five -Outpost scenarios across all three models, not fifteen-plus scenarios on the -weak model, and the two populations answer different questions. They were -conflated once in conversation into a claim that the delta's *sign* had flipped, -which no run supports. State the scenario set and the model with any skills -delta, or it will be compared against a number measuring something else. +Two things that output will not tell you. The denominators are the filter: a row +reading `2/2` is a partial re-run of a couple of cells, not a suite delta. And +`runId` only exists from 1 September, so earlier executions fall back to their +date and two on the same day merge — which is another reason 13 August needs +reading by cause rather than by row. + +**State the scenario set and the model with any skills delta**, or it will be +compared against a number measuring something else. Loop 2's corrected +measurement was `+skills` 12/12 against `-no-skills` 10/12, but that is five +Outpost scenarios across all three models rather than nineteen scenarios on the +weak model, and the two populations answer different questions. That they now +point the same way is not a reason to merge them. **Product findings come from runs, not speculation.** Several concern `hookdeck listen`: it crashes without a TTY unless given `--output compact`; a