Skip to content

fix(llm): consolidate fallback into one physical attempt budget - #902

Merged
drewstone merged 9 commits into
mainfrom
fix/shared-llm-attempt-budget
Sep 30, 2026
Merged

drewstone merged 9 commits into
mainfrom
fix/shared-llm-attempt-budget

Conversation

@drewstone

@drewstone drewstone commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Outcome

One existing retry loop owns raw and structured model requests. Schema fallback no longer starts a second call with a fresh attempt allowance, deadline origin, or raw-provider attempt sequence.

  • maximumAttempts counts physical HTTP requests, including transient, schema and temperature fallback. One means one request.
  • Preserve temperature fallback, raw-call no-schema-negotiation behavior, explicitly selected JSON-object mode, operation identity and result validation.
  • maximumChargeForLlmRequest reserves that same allowance instead of two full batches for schema-capable calls. The original schema still bounds input; output limits remain enforced per physical request.
  • Remove the nested retry owner rather than adding an executor, policy class, or product wrapper.
  • Keep older products' fetch guards until an actual corrected library release is published, installed and verified.

Final verification

Head: 64aeaf8683fff99bd9c99905d20a3ba1dea22717.
Original base: a6ed111e533b1af74792a9c727e2bced7065e8b6.

Full exact-head CI passed: https://github.com/tangle-network/agent-eval/actions/runs/36649148569

Inspected successful steps: dependency installation, Biome, source/scripts/examples typechecks, full tests, build/OpenAPI, packed exports, Python client/distributions, official TypeScript optimizer integration, GEPA and DSPy compatibility. Both main CI and the separate published-GEPA job passed.

Local proof: 96 focused checks passed, including seven real-loopback HTTP cases and the maintained client/retry/raw-capture/analyst suites. Five of the seven new HTTP cases fail against the exact unchanged production baseline, then pass after the fix. Native Node fetch and the public client execute; the external response is controlled by a real local HTTP server.

Follow-through from CI

  • Corrected a receipt fixture that claimed $0.25 despite declaring one physical request whose bound was $0.167535. It now returns and asserts $0.01; all three findings and single-receipt conservation assertions remain. Cancellation tests retain their original paid receipts.
  • Updated only the CURRENT benchmark implementation digest to the value computed by the repository checker. Historical published-evidence digests and dependency identity are unchanged; no old accuracy claim is reassigned to the new implementation.
  • Added .gitattributes for LF TypeScript/Markdown checkouts after reproducing the formatter failure with core.autocrlf=true. No formatter gate was removed or relaxed.
  • Two existing reservation expectations change because the phantom second request batch is gone, not because a limit was increased.

Local execution used a restored dependency tree; the final repository CI above is separate stronger evidence. No test was skipped, no coverage threshold lowered, and no signoff was fabricated.

Reproduce

pnpm install --frozen-lockfile
pnpm exec vitest run src/llm-client.test.ts tests/llm-transient-status.test.ts tests/llm-raw-capture.test.ts tests/llm-physical-attempts.test.ts src/analyst/adapters.test.ts
pnpm typecheck
pnpm build
pnpm verify:package

Boundaries

This fixes attempt ownership, not provider-side deduplication, automatic result replay, aggregation of all failed paid-attempt receipts into the terminal response, or every single-attempt body-read deadline. Existing caller accounting still owns observations. No dependency, package version, credentials, production deployment, funded inference or merge changed.

The temporary exact-source importer was isolated on a delivery branch and removed after use; it is absent from the feature diff. It executed no candidate application code and wrote no success status.

Refs #768; complements tangle-network/agent-dev-container#8533 without a cross-repository release dependency.

tangletools
tangletools previously approved these changes Sep 29, 2026

@tangletools tangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — 05ccfc4f

Blanket team auto-approval is intentional. The merge gates are CI and codex-p1.
No automated review runs on this PR. This approval rests on the rule above alone.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-09-29T23:41:56Z

tangletools
tangletools previously approved these changes Sep 30, 2026

@tangletools tangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — 534ad8ae

Blanket team auto-approval is intentional. The merge gates are CI and codex-p1.
No automated review runs on this PR. This approval rests on the rule above alone.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-09-30T00:00:56Z

tangletools
tangletools previously approved these changes Sep 30, 2026

@tangletools tangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — c33ffad1

Blanket team auto-approval is intentional. The merge gates are CI and codex-p1.
No automated review runs on this PR. This approval rests on the rule above alone.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-09-30T00:11:56Z

tangletools
tangletools previously approved these changes Sep 30, 2026

@tangletools tangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — 64aeaf86

Blanket team auto-approval is intentional. The merge gates are CI and codex-p1.
No automated review runs on this PR. This approval rests on the rule above alone.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-09-30T00:12:56Z

tangletools
tangletools previously approved these changes Sep 30, 2026

@tangletools tangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — ce18ea63

Blanket team auto-approval is intentional. The merge gates are CI and codex-p1.
No automated review runs on this PR. This approval rests on the rule above alone.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-09-30T05:56:32Z

tangletools
tangletools previously approved these changes Sep 30, 2026

@tangletools tangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — 6135051e

Blanket team auto-approval is intentional. The merge gates are CI and codex-p1.
No automated review runs on this PR. This approval rests on the rule above alone.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-09-30T05:59:35Z

@tangletools tangletools left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Auto-approved PR — 41afe73e

Blanket team auto-approval is intentional. The merge gates are CI and codex-p1.
No automated review runs on this PR. This approval rests on the rule above alone.

tangletools · auto-approval · reason: blanket_auto_approve · 2026-09-30T06:00:32Z

@drewstone
drewstone merged commit 443abdf into main Sep 30, 2026
3 of 4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants