feat: take the context window from the backend instead of guessing it - #342
Merged
Merged
Conversation
codeoid inferred every model's context window from a substring table over
model ids. That is wrong the moment a model ships: `claude-opus-5-5` inferred
to the 200k fallback while every turn result reported 1,000,000, and everything
sized against it was off by 5x — the percent-of-window display, the fork seed
budget, and the auto-rotate occupancy that decides when a session rolls, so
auto-rotate fired at roughly a fifth of real capacity.
The backends already knew. What they report is now cached per
(provider, model) and persisted, so the number survives a restart and the first
session to learn a model's window teaches every later one. A window is a
property of the model, not of whoever happened to run the turn.
Audited across every backend codeoid drives, because they do not agree on how
they publish it. Two ingresses were needed to cover them:
claude per turn result.modelUsage[model].contextWindow
codex per turn thread/tokenUsage/updated -> tokenUsage.modelContextWindow
qwen per MODEL its catalog's contextWindowSize — before any turn runs
gemini nothing direct API; the response carries no window
openai nothing same
pi nothing same
acp nothing gemini-cli over ACP publishes no limits
Two of those were already on the wire and being discarded. qwen's is the best
signal of the three — it is known before the first turn — and
`normalizeModelCatalog` projected the catalog down to id/label/description and
dropped it. codex's arrived on `thread/tokenUsage/updated`, a notification the
provider already handled, reading two of its three fields; the codex binary's
own serde descriptors name the third ("struct ThreadTokenUsage with 3
elements" / `modelContextWindow`).
The static tables stay, demoted to what they should always have been: a
bootstrap. Nothing can know the window before a turn completes on the two
per-turn backends, and four backends report none at all, so an inferred floor
is structurally required — it is now the last resort rather than the only
answer. It stays a positive number so a silent backend never divides the
percent-of-window by zero, and the reported value is sticky per session: a
provider that reports on turn 1 and not on turn 5 has not said the window
changed, and flapping back to inference would make the number oscillate
between correct and wrong.
Fixed at the source too. The Claude provider derived the turn's model from
`Object.keys(modelUsage)[0]`, but that map is keyed by every model the turn
touched and ordered by insertion — a Haiku side-call (title generation)
routinely sits in front of the model that did the work. Measured on the live
backend: a single `opus` turn yields
`{claude-haiku-4-5-...: {...}, claude-opus-5-5: {...}}`. So an Opus turn was
labelled Haiku in canonical history, and reading a window off that first entry
would have reported 200k for a 1M turn — the same bug, at the point meant to
cure it. The SDK names the primary model on its `init` message; that is used
instead.
Every ingress is mutation-tested: dropping the session's preference, codex's
field, or qwen's field each turns a test red.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…d-context-window # Conflicts: # CHANGELOG.md
jalbrethsen-highflame
approved these changes
Sep 24, 2026
… running
A deep audit of the first commit confirmed 23 findings. Most share a root
cause: the window was resolved in five places from five different mixes of
sources, and the backend-stated number was a bare value that outlived the
model it described.
The worst was a regression. After `/provider codex` from an Opus session the
sticky 1M sized codex's history seed at 2,450,000 chars instead of 627,200, so
a long transcript went whole into a 272k window with no truncation notice, and
the display kept showing 1M on a backend that never reports.
Fixed structurally rather than site by site:
- One resolver, `Session#contextWindow`, for the display, the occupancy caps,
auto-rotate and the rotate message; the seed and its truncation notice share
its stated-window lookup, so the notice can no longer cite a different
window than the one that sized the budget.
- The stated window is an observation of (provider, model, window), cleared
wherever the model or provider changes and read only for its own provider.
- A fork on the same (provider, model) inherits its parent's observation, so
it seeds against the real window instead of the 200k floor.
- The display is computed at read time, so it is set with the memory engine
off (it was never emitted, and the web UI divided by 200k).
- Auto-rotate divided by a fixed 1M. Below ~970k the 0.97 hard ceiling could
never fire, so 200k and 272k sessions never rotated. It now rotates against
the running model's window.
The remembered windows are scoped by account, project and workdir. The Claude
CLI derives the number from settings a workdir can override —
CLAUDE_CODE_MAX_CONTEXT_TOKENS is unclamped — so an unscoped cache let one
workdir's 4M size another tenant's fork seeds. Catalog-published windows are
read from the catalog itself instead of being copied into an upsert-only table
that kept them after the backend stopped publishing.
Backends:
- codex: reset the window per turn. The instance outlives /model, so a turn
interrupted before its first usage notice reported the previous model's
window under the new model's id.
- pi: states its window on `get_session_stats` (made every turn) and on its
model objects. Both were dropped; the first commit had listed pi as silent.
- qwen: `unionCatalogs` dropped the window whenever the gateway also listed the
model. It now carries it across whichever entry wins the label.
- claude: follow `model_fallback`, so a turn the fallback model served is
labelled with that model and its window, not the primary's.
- Placeholder model ids ("unknown", codex's "codex", pi's "pi-default") are
never cache keys.
Also: the TUI now reads the daemon's window like the web UI; `maxOutputTokens`
is gone (carried and persisted, never read); the catalog entry has one shared
type at every hop so a re-projection can't silently drop the window; the
misplaced `_cacheModels` JSDoc is reattached; eight pasted hook pairs are one
spread; stale comments and the CHANGELOG are corrected — auto-rotate was never
sized from the table, it was pinned to 1M, which failed the other way.
Every fix is guarded: the switch, rotation, memory-off display, model-switch
clear, codex reset and pi window each turn a test red when reverted.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #341, which patched a stale window table. This removes the reason the table could go stale.
The problem
codeoid inferred every context window from a substring table over model ids. That's wrong the moment a model ships —
claude-opus-5-5inferred to the 200k fallback while every turn reported 1,000,000 — and everything sized against it was off by 5×: the percent-of-window display, the fork seed budget, and the auto-rotate occupancy that decides when a session rolls. Auto-rotate was firing at roughly a fifth of real capacity.The backends already knew. codeoid was discarding it.
Audited across every backend, because they disagree on how they publish it
You asked me to make sure this works for all of them. It needed two ingresses, not one:
result.modelUsage[model].contextWindowthread/tokenUsage/updated→tokenUsage.modelContextWindowcontextWindowSizeTwo of the three were already arriving and being thrown away:
normalizeModelCatalogprojected the catalog down to id/label/description and dropped the window.struct ThreadTokenUsage with 3 elements/modelContextWindow.So the design is: two ingresses, one cache. Catalog (pre-turn, when a backend publishes it) and turn result (post-turn). Both land in a
(provider, model)-keyed cache, persisted — so it survives a restart, and the first session to learn a model's window teaches every later one. A window is a property of the model, not of whoever ran the turn.The tables stay — demoted, on purpose
This can't go to zero hardcoding, and I'd rather say so than imply otherwise:
ModelInfothere carries none).So an inferred floor is structurally required. It's now the last resort instead of the only answer, it stays a positive number so a silent backend never divides percent-of-window by zero, and the reported value is sticky per session: a provider that reports on turn 1 and not turn 5 hasn't said the window changed, and flapping back to inference would oscillate the number between correct and wrong.
Fixed at the source too
The Claude provider derived the turn's model from
Object.keys(modelUsage)[0]. That map is keyed by every model a turn touched, ordered by insertion — and a Haiku side-call (title generation) routinely sits in front of the model that did the work. Measured live, oneopusturn:{ "claude-haiku-4-5-20251001": {...}, "claude-opus-5-5": {...} }So an Opus turn was labelled Haiku in canonical history — and reading a window off that first entry would have reported 200k for a 1M turn, i.e. the same bug at the point meant to cure it. The SDK names the primary model on its
initmessage; that's used now.Verification
Every ingress is mutation-tested — dropping the session's preference, codex's field, or qwen's field each turns a test red:
session-reported-window.test.ts(new): reported beats inference; inference still answers pre-turn and when a backend reports nothing; sticky across an omitting turn; follows a changed window; hands limits up for caching; per-provider ingress table including the four silent backends.models.test.ts: cache keyed per (provider, model), survives a restart, refuses 0/negative/unknown, latest report wins.provider-codex.test.ts+ the fake-codex fixture (which was missing the third field entirely).provider-qwen.test.ts: the real 0.1.8 catalog shape now carries the window.context-windows.test.ts: reported outranks the per-model table and the per-provider default; ignored when absent/non-positive.bun run typecheck·bun run lint·bun test→ 2613 pass, 19 skip, 0 fail. Web: 553 pass.Possible next step, not done here
Gemini's
models.getexposesinputTokenLimit, so gemini could join the catalog ingress with an extra API call at startup. I left it out — it's a new network call on a path that currently makes none, and the four silent backends are correct-by-fallback today.🤖 Generated with Claude Code