Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions .release-notes/router-auto-evidence.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
type: minor
---
Router tool executors support versioned inline Auto policies without weakening concrete model checks.
Runtime checks every declared strategy and assessment against the allowed models before dispatch.
Buffered HTTP response receipts preserve complete routing metadata and failed responses in native transcripts.
An awaited provider response observer supports durable capture during execution.
Aggregate Auto usage remains unknown because its response reports final-generation tokens only.
5 changes: 3 additions & 2 deletions api-surface.json
Original file line number Diff line number Diff line change
Expand Up @@ -1234,8 +1234,9 @@
"RootProviderModelEvidence": "type 062cb80ee9d2",
"RootSignal": "type b819c5206299",
"RootStreamReceipt": "type e19150df084b",
"RouterResponseReceipt": "type 4abf1e70a35a",
"RouterSeam": "type 3f4c9e0a841e",
"RouterToolsSeam": "type 44e34624f17c",
"RouterToolsSeam": "type c5dc09516fca",
"RouterTransportConfig": "type d14ea3416352",
"RunAgentRoundsOptions": "type 4caa711f9f2e",
"RunAgenticOptions": "type 346e500a1693",
Expand Down Expand Up @@ -2111,7 +2112,7 @@
"./testing": {
"AgentProfileImprovementFixture": "value 8a79d0f5c646",
"AgentProfileImprovementProposalFixture": "value 95f42c91bc08",
"DriverAgentOptions": "type 52612582f058",
"DriverAgentOptions": "type 605daae26739",
"RunGraphTestOptions": "type c5b5730ba4f0",
"SuperviseTestOptions": "type 1fd4a3f892e4",
"SupervisorAgentTestDeps": "type f39aa3b16149",
Expand Down
3 changes: 2 additions & 1 deletion docs/api/primitive-catalog.md
Original file line number Diff line number Diff line change
Expand Up @@ -449,7 +449,7 @@ Import from `@tangle-network/agent-runtime/intelligence` — 167 exports.

### Execution kernel — recursive atom, supervision, executors, round-synchronous loop

Import from `@tangle-network/agent-runtime/kernel` — 1043 exports.
Import from `@tangle-network/agent-runtime/kernel` — 1044 exports.

| Symbol | Kind | Summary |
|---|---|---|
Expand Down Expand Up @@ -1069,6 +1069,7 @@ Import from `@tangle-network/agent-runtime/kernel` — 1043 exports.
| `RetainedRunStartMaterial` | interface | Environment, turn, and optional identity needed to replay one retained start. |
| `RootHandle` | interface | Live root handle — a chat/pi-viz client uses it to inspect and control one root run. |
| `RootStreamReceipt` | interface | The root manager's retained provider stream: `<runDir>/root-stream.jsonl`, one line per |
| `RouterResponseReceipt` | interface | Exact buffered HTTP evidence. Authentication and cookie headers are excluded. |
| `RouterSeam` | interface | Router/inline transport seam. The profile owns model, prompt, and generation behavior. |
| `RouterToolsSeam` | interface | Router seam WITH tool use — the tool-using router backend. Same direct |
| `RouterTransportConfig` | interface | Connection details for Runtime's Router-backed executors. |
Expand Down
70 changes: 69 additions & 1 deletion docs/api/runtime.md
Original file line number Diff line number Diff line change
Expand Up @@ -1417,7 +1417,7 @@ One flattened node with the journal tree that owns its records.

###### Inherited from

[`NodeSnapshot`](#nodesnapshot).[`status`](#status-20)
[`NodeSnapshot`](#nodesnapshot).[`status`](#status-21)

##### runtime

Expand Down Expand Up @@ -9591,6 +9591,58 @@ Injectable OpenAI-compatible transport. Optional usage.resources carries measure

***

### RouterResponseReceipt

Exact buffered HTTP evidence. Authentication and cookie headers are excluded.

#### Properties

##### endpoint

> `readonly` **endpoint**: `string`

##### callId

> `readonly` **callId**: `string`

##### attempt

> `readonly` **attempt**: `number`

##### startedAt

> `readonly` **startedAt**: `string`

##### endedAt

> `readonly` **endedAt**: `string`

##### status

> `readonly` **status**: `number` \| `null`

##### error?

> `readonly` `optional` **error?**: `object`

###### name

> `readonly` **name**: `string`

###### message

> `readonly` **message**: `string`

##### headers

> `readonly` **headers**: `Readonly`\<`Record`\<`string`, `string`\>\>

##### body

> `readonly` **body**: `string` \| `null`

***

### ToolSpec

#### Properties
Expand Down Expand Up @@ -20842,6 +20894,22 @@ surfaces (e.g. a gym keyed by task) can dispatch correctly.

Exact conversation to continue. Runtime validates its system message against the profile.

##### onProviderResponse?

> `optional` **onProviderResponse?**: (`receipt`) => `void` \| `Promise`\<`void`\>

Persist each buffered HTTP attempt, including failed responses, before the next turn.

###### Parameters

###### receipt

[`RouterResponseReceipt`](#routerresponsereceipt)

###### Returns

`void` \| `Promise`\<`void`\>

##### onMessages?

> `optional` **onMessages?**: (`messages`) => `void` \| `Promise`\<`void`\>
Expand Down
26 changes: 26 additions & 0 deletions docs/canonical-api.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,32 @@ Run pnpm docs:freshness after editing this file. -->
>
> **Read this before writing any orchestration, optimization, or measurement code in this repo.** If you are about to write a persona⟷agent conversation runner, a "skill optimizer", a "profile-seam", a depth-vs-breadth A/B harness, a bootstrap loop, or a `new Sandbox(...)` + stream + read dance: **stop**, it already exists, and a parallel copy will silently break one of the guarantees the existing primitives enforce: equal compute per compared arm ("equal-k"), the attempt-picker never being the grader ("selector≠judge"), complete usage capture, or eval running the same code path as production.

## Router Auto execution and evidence

Use `createExecutor({ backend: 'router-tools', ... })` with an exact profile.
Declare `model.default: 'tangle/auto'` and an inline policy in `model.metadata.extraBody.auto`.
Use the Router Auto base URL and buffered transport.
Runtime checks strategy and assessment models against `allowedModels` before dispatch.
It checks each reported child model against the declared policy.
It checks policy ID and revision; Router owns normalization and the policy digest.
Consumers must compare that digest with their frozen Router policy before using the result.
Preset references require resolution into an inline policy before execution.

`onProviderResponse` receives each buffered HTTP attempt before response decoding.
The receipt preserves its full body, endpoint, call identity, times, status, and response headers.
Authentication and cookie headers are excluded.
Network failures have no HTTP status and retain their error.
Native transcripts preserve these receipts outside the model conversation.
Final outputs also retain received HTTP receipts.
A failed attempt still retains its transcript.
Observer failures stop the executor and do not authorize another paid request.

Auto response usage covers the final generation only.
Runtime marks aggregate token and dollar usage unknown.
Any observed final-generation dollar subtotal is a lower bound, not a complete Auto total.
Use the returned physical generation identifiers to join authoritative billing records.
A receipt header cost does not prove complete campaign billing.

## 1. Mental model: the spine

> **Legend**: five terms the rest of this doc leans on, in plain terms:
Expand Down
2 changes: 1 addition & 1 deletion src/runtime/index.ts
Original file line number Diff line number Diff line change
Expand Up @@ -481,7 +481,7 @@ export {
// Router requests are an internal transport adapter. Public execution always enters through an
// exact AgentProfile (`createExecutor` + `streamAgentTurn`); callers may configure only the
// endpoint/auth transport used by that path.
export type { RouterTransportConfig } from './router-client'
export type { RouterResponseReceipt, RouterTransportConfig } from './router-client'
export {
type BenchmarkCell,
type BenchmarkConfig,
Expand Down
103 changes: 103 additions & 0 deletions src/runtime/router-auto-evidence.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,103 @@
import { ValidationError } from '../errors'
import { observedModelMatchesDeclared } from './model-identity'

function object(value: unknown): Record<string, unknown> {
if (value === null || typeof value !== 'object' || Array.isArray(value)) {
throw new ValidationError('Router Auto requires an inline policy and routing receipt')
}
return value as Record<string, unknown>
}

/** Router owns policy parsing. Runtime requires visible model authority before dispatch. */
export function assertRouterAutoDeclaration(value: unknown, stream: boolean | undefined): void {
const policy = object(value)
if (stream === true) throw new ValidationError('Router Auto requires buffered transport')
if (
policy.version !== 1 ||
typeof policy.id !== 'string' ||
typeof policy.revision !== 'string' ||
!Array.isArray(policy.strategies) ||
policy.strategies.length === 0
) {
throw new ValidationError(
'Router Auto requires a versioned inline policy with concrete strategies',
)
}
for (const raw of policy.strategies) {
const strategy = object(raw)
if (
typeof strategy.id !== 'string' ||
typeof strategy.model !== 'string' ||
strategy.model === 'tangle/auto' ||
strategy.model.length === 0
) {
throw new ValidationError('Router Auto strategy model authority is missing')
}
}
for (const phase of ['selector', 'review']) {
if (policy[phase] !== undefined && typeof object(policy[phase]).model !== 'string') {
throw new ValidationError('Router Auto assessment model authority is missing')
}
}
}

/** An alias does not waive identity checks. Every reported child must match its declared phase. */
export function assertRouterAutoResponse(
policyValue: unknown,
response: unknown,
served: string | undefined,
): void {
const policy = object(policyValue)
const receipt = object(object(response).tangle_auto)
const decision = object(receipt.decision)
const strategies = policy.strategies as unknown[]
const selected = strategies.map(object).find((strategy) => strategy.id === decision.strategyId)
if (
decision.requestedModel !== 'tangle/auto' ||
decision.policyId !== policy.id ||
decision.policyRevision !== policy.revision ||
typeof decision.policyDigest !== 'string' ||
!selected ||
decision.selectedModel !== selected.model ||
served === undefined ||
!observedModelMatchesDeclared(served, String(selected.model))
) {
throw new ValidationError(
'Router Auto response does not match AgentProfile policy/model authority',
)
}
if (!Array.isArray(receipt.calls) || receipt.calls.length === 0) {
throw new ValidationError('Router Auto response is missing physical call receipts')
}
for (const raw of receipt.calls) {
const call = object(raw)
const declared =
call.phase === 'selector'
? object(policy.selector).model
: call.phase === 'review'
? object(policy.review).model
: call.phase === 'initial' || call.phase === 'escalation'
? strategies.map(object).map((strategy) => strategy.model)
: []
const allowed = Array.isArray(declared) ? declared : [declared]
if (
typeof call.model !== 'string' ||
!allowed.some(
(model) =>
typeof model === 'string' && observedModelMatchesDeclared(call.model as string, model),
)
)
throw new ValidationError('Router Auto physical call exceeds AgentProfile model authority')
}
}

/** The concrete authority declared by an inline Auto policy, including paid assessments. */
export function routerAutoDeclaredModels(value: unknown): readonly string[] {
assertRouterAutoDeclaration(value, undefined)
const policy = object(value)
const models = (policy.strategies as unknown[]).map((raw) => String(object(raw).model))
for (const phase of ['selector', 'review']) {
if (policy[phase] !== undefined) models.push(String(object(policy[phase]).model))
}
return [...new Set(models)]
}
Loading
Loading