Conversation
The Status section read "the weak model is two worse with skills than without" and offered a recompute snippet as the corrective. Both were wrong, and the snippet is why: it iterates snapshot files rather than executions, so the 17 August pair appears in six of them and reads -2 in all six. That is where "eight of the nine later runs read -2" came from. Keyed per execution, the five full-suite pairwise runs read -3, -2, -1, +1, +4. #2 closed on 22 September on that basis: pooled across the three September executions the weak model is +skills 48/57 against 44/57, all of the gain on Outpost, Event Gateway level at 34/42 in both arms. Also corrects two other figures in the same section against results/latest.json: seven of nineteen scenarios are failed by at least one experiment, not nine, and the frontier deltas in that file are not deltas at all -- it is a merged snapshot whose -no-skills arms ran on 1 September and whose +skills arms ran on 14 September. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The task
Triage every open issue on this repository: check each is clear, check none has
already been addressed, and propose a roadmap of refine / implement / close. The
triage itself is reported back to the requester — per
AGENTS.md, work lives inGitHub Issues and order lives in #24, not in a markdown file that goes stale.
This PR is the one thing the triage found wrong in the repository rather than
on the board.
What changed
The Status section of
AGENTS.md, three claims and one recipe.1. The recompute snippet counted snapshot files, not executions. A snapshot
republishes the cells a run did not re-measure (#60), so iterating
results/runs/*.jsoncounts one execution many times. The 17 August pair appearsin six snapshot files and reads −2 in all six — which is exactly where the
paragraph's "eight of the nine later runs read −2" came from. The replacement
keys each cell by
runId(falling back toranAt) before counting.2. "The weak model is two worse with skills than without" is not supported by
any run since August. Keyed per execution, the five full-suite pairwise runs of
that pair read:
Pooled over the three September executions:
+skills48/57 against 44/57, withall of the gain on Outpost (14/15 against 10/15) and Event Gateway level at 34/42
in both arms. Those are #2's closing figures, reproduced here from
results/runs/rather than quoted. #2 was closed as not planned on 22 Septemberon that basis, and this file kept the old sign for a further two days — and the
old figure for six weeks.
3. Two smaller figures in the same section. "Nine of nineteen scenarios
discriminate" is seven on the published snapshot. And the Claude and GPT-5.6
skills deltas readable from
results/latest.jsonare not deltas: it is a mergedsnapshot whose frontier
-no-skillsarms were executed on 1 September and whose+skillstwins on 14 September, so subtracting them compares two instruments afortnight apart. Only the weak pair runs weekly, so only the weak pair is a clean
comparison in that file.
The methodology rule that followed — state the scenario set and the model with
any delta — is kept. What is dropped is its worked claim that no run supports the
sign flipping, because three now do.
What I verified, and how
Every figure above was recomputed from
results/runs/andresults/latest.jsonin this worktree, not taken from an issue. The snippet as it now appears in the
file was extracted from the markdown and run verbatim; the output in this
description is that run.
Nothing was already failing.
pnpm installneeds/opt/homebrew/binahead ofthe asdf shims on
PATH— the repo has no.tool-versionsand the shim has nopnpmversion set. Unrelated to this change; already noted on #82.No tests
The change is prose and a shell snippet in
AGENTS.md. Nothing here is importedby anything. The snippet was verified by executing it out of the file, which is
the only check available to it.
Deliberately not done
proposals — which issues to close, which to refine, and the order to work in —
went back to the requester. Closing is a human decision and Roadmap: the order we are working in, and why #24 is the
requester's to update.
"+1 for Claude, 0 for GPT-5.6 and −2 for the deliberately weak model". Correcting
it is an issue edit, not a diff, and is in the proposal.
.plans/delivery-plan.mdwas not audited. It was last brought back to trueon 29 August (Bring the plan's status, counts and costs back to what is true #69) and may carry the same figure; checking it is a separate pass
and would have widened this diff past its description.
1 September and +4 on 14 September to make a different point — that movement is
not a reason to release — and it is accurate. It omits 7 September's +1, which
strengthens rather than weakens it.
Observations, not changed here
at
fb16bab.hookdeck ciappears in zero of the three pinned skills, soDoes documenting hookdeck ci stop agents giving up on CLI authentication? #27 — a mapping issue open since 19 August, waiting on a run to measure a merged
skills change — cannot be measured by any run we have done or will do until the
pin moves. The skills submodule pin drifts silently, and moving it will change results with no record #26 is the pin policy and the two are one piece of work.
usage.costUsdis now on 0 of 114 published rows, down from the 4 of 114that Review and merge or drop the Codex cost-reporting branch #6 records. The cost multiplier
kin the What this costs section isless derivable than when it was written, not more.
results/latest.jsonis 10 days old and the 21 September weekly publishednothing (One errored cell stops a whole run from publishing #80). The release pinned to the page is v0.4.0, from 2 September.
🤖 Generated with Claude Code