From 1b5a02304cc59ce3683f6ce46dc2bf375392f3cb Mon Sep 17 00:00:00 2001 From: Tiberiu Socaci Date: Mon, 28 Sep 2026 15:39:51 +0300 Subject: [PATCH 01/16] docs(release): the updater rebuilds the 0.6.0 image; no manual build needed Co-Authored-By: Claude Fable 5.1 Signed-off-by: Tiberiu Socaci --- docs/RELEASE-CHECKLIST.md | 5 +++-- 1 file changed, 3 insertions(+), 2 deletions(-) diff --git a/docs/RELEASE-CHECKLIST.md b/docs/RELEASE-CHECKLIST.md index 39386486..eeb8c48e 100644 --- a/docs/RELEASE-CHECKLIST.md +++ b/docs/RELEASE-CHECKLIST.md @@ -16,8 +16,9 @@ > checkbox and optional Allowed domains, non-admin turns are read-only while an Admin channel > mounts the operator home, Codex Read mode runs commands again, the Claude model pickers show exact > versions, and Codex usage metrics are switched off on every spawn. The image is spec 1.6.1 -> (bubblewrap, per-channel `/tmp` volumes); every other host must run `npm run build:image` before -> restarting. Live on Xavier before the cut (Claude and Codex through the QA actors, recorded in the +> (bubblewrap, per-channel `/tmp` volumes); the updater (Settings → Update or `scripts/update.sh`) +> rebuilds the default image automatically — `npm run build:image` by hand only if it reports the +> image build failed or the host uses a custom image. Live on Xavier before the cut (Claude and Codex through the QA actors, recorded in the > private QA base): EGR-01/03/04/05/07/09/10, CDX-01..04/08, EN-01 and EN-09 on Codex, CTR-30 on both > engines, SSHP-01..05 from the gateway host, and the hidden-variable first-use approval with the > post-approval swap. Still open for the next release: SSHP-06/07 (need a desktop VS Code), the Slack From 991272b951381fe0e78d597315e47cc2780ce4a4 Mon Sep 17 00:00:00 2001 From: Tiberiu Socaci Date: Mon, 28 Sep 2026 15:44:59 +0300 Subject: [PATCH 02/16] fix(update): a provider usage limit counts as reachable in the engine smoke The update's engine smoke refused an update on Atlas because the Claude plan had hit its weekly limit. That answer proves the CLI started in the new container, authenticated and reached its provider, which is what the smoke exists to prove. It now counts as reachable, with a "reachable but usage-limited" note, before the update and after the restart. Real failures (a bad login, a missing CLI, a wrong answer) still refuse the update or roll it back. Co-Authored-By: Claude Fable 5.1 Signed-off-by: Tiberiu Socaci --- CHANGELOG.md | 4 ++++ FEATURES.md | 8 +++++++- TEST-PLAN.md | 6 ++++++ src/gateway/update-smoke.js | 18 ++++++++++++++++++ test/update-smoke.test.js | 36 ++++++++++++++++++++++++++++++++++++ 5 files changed, 71 insertions(+), 1 deletion(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index b1fbced5..825cf8ea 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -16,6 +16,10 @@ product overview. > | Makeitfuture Sustainable Use License 1.1 | 2026-08-20 | never published | > | Makeitfuture Sustainable Use License 1.0 | 2026-08-06 | never published | +## Unreleased + +- **A provider usage limit no longer blocks an update.** Before and after updating, the updater sends each engine a tiny test prompt inside a throwaway container. A reply like "You've hit your weekly limit" used to refuse the update, even though it proves the engine starts, logs in and reaches its provider. It now counts as reachable. A broken login or engine still stops the update. + ## 0.6.0 — 2026-09-28 - **Claude model choices now show exact versions.** The `/model` picker and admin selectors list diff --git a/FEATURES.md b/FEATURES.md index 3203cb97..e9c72885 100644 --- a/FEATURES.md +++ b/FEATURES.md @@ -310,7 +310,13 @@ A categorized catalog of what's shipped. Cross-linked to `TEST-PLAN.md` checks. Its script refuses non-hosted or occupied hosts. A separate KVM guest workflow exercises an actual OS reboot, service autostart and persistent database/container-volume fixtures. Authenticated engine update smoke and conversation/session acceptance remain separate live - gates. → TEST-PLAN: Disposable Linux lifecycle workflow. + gates. → TEST-PLAN: Disposable Linux lifecycle workflow. The update's engine smoke (a + throwaway container, one fixed prompt per configured engine, before the update and after the + restart) counts a provider **usage limit** (weekly/session plan cap, spent credits, rate limit) + as REACHABLE, noted "reachable but usage-limited": the CLI started in the new container, + authenticated and reached its provider, which is what the smoke proves (2026-09-28, an update on + Atlas was refused over a Claude weekly limit). A bad login, a missing CLI or a wrong answer still + refuses the update or rolls it back. - **Service-account image provisioning** uses the same explicit environment as the systemd daemon so operator XDG/container storage settings cannot redirect a fresh build into another user's private Podman store, and runs from the service-owned checkout so an operator-private diff --git a/TEST-PLAN.md b/TEST-PLAN.md index f4263541..b340e5bf 100644 --- a/TEST-PLAN.md +++ b/TEST-PLAN.md @@ -7869,6 +7869,12 @@ the suite runs as an enterprise deployment because it holds a license it actuall ## Container update verification and recovery +- [x] Unit — the update engine smoke counts a provider usage limit (thrown or returned: Claude's + weekly/session limit, Codex's usage limit, a rate limit) as reachable with a + "reachable but usage-limited" note, both before the update and for an engine required after + the restart; an authentication failure, a missing CLI or an unexpected answer still fails it + (`test/update-smoke.test.js`). + - Both Claude and Codex fixtures: configure each login, invoke the internal authenticated update smoke route, require exact CG_UPDATE_SMOKE_OK responses per engine. Verify isolated target, no bypass/MCP injection, no host engine child, and removal of the ephemeral container, HOME volume and work folders. Invalid configured credentials must fail; absent credentials must be explicitly skipped; zero probes fails. - Image recovery: build old pins, change the desired Codex pin without changing imageSpecVersion, make the first build fail, then run Update again on the same checkout revision. Require retry and matching built/desired source digest; a build exiting zero with stale labels fails verification. Existing channels retain HOME and adopt the new image when idle. - Admin container status: require actual built and desired Claude/Codex versions, rebuild-needed status, and count of containers awaiting adoption. A custom image reference must never report a successful default-image rebuild. diff --git a/src/gateway/update-smoke.js b/src/gateway/update-smoke.js index 9ce1b203..caf743fb 100644 --- a/src/gateway/update-smoke.js +++ b/src/gateway/update-smoke.js @@ -71,6 +71,16 @@ export async function prepareUpdateSmokeFolder({ }; } +// A provider USAGE LIMIT (a weekly/session plan cap, spent credits, a rate limit) is not a broken +// gateway: to get that answer the engine CLI started inside the new container, authenticated with +// the configured login and reached its provider — everything this probe exists to prove. Refusing +// an update over it (live on Atlas, 2026-09-28: "claude: Claude usage limit reached: You've hit your +// weekly limit") would hold every host hostage to a quota. Such an engine counts as reachable. +const USAGE_LIMIT_RE = /usage limit|(?:hit|reached|exceeded)[^.\n]{0,40}(?:session|usage|weekly|monthly|daily|account|spend|credit|rate|quota|token)\s+limit|rate[_ -]?limit|quota exceeded|out of credits?|purchase more credits|insufficient (?:credits?|quota|balance|funds)|claude\.ai\/settings\/usage|codex\/settings\/usage/i; +export function usageLimitedSmoke(text) { + return USAGE_LIMIT_RE.test(String(text ?? "")); +} + export function validSmokeResponse(content) { return String(content ?? "").trim() === RESPONSE; } @@ -122,9 +132,17 @@ export async function runUpdateSmoke({ timeoutMs, }); const ok = validSmokeResponse(result?.content); + if (!ok && usageLimitedSmoke(result?.content)) { + results.push({ engine: engine.id, ok: true, usageLimited: true, durationMs: Date.now() - engineStart, note: `reachable but usage-limited: ${String(result?.content || "").split(/\r?\n/, 1)[0].slice(0, 160)}` }); + continue; + } results.push({ engine: engine.id, ok, durationMs: Date.now() - engineStart, ...(!ok ? { error: "Engine smoke probe returned an unexpected response." } : {}) }); } catch (error) { + if (usageLimitedSmoke(error?.message)) { + results.push({ engine: engine.id, ok: true, usageLimited: true, durationMs: Date.now() - engineStart, note: `reachable but usage-limited: ${safeError(error)}` }); + continue; + } results.push({ engine: engine.id, ok: false, durationMs: Date.now() - engineStart, error: safeError(error) }); } } diff --git a/test/update-smoke.test.js b/test/update-smoke.test.js index 33442b04..e8a56980 100644 --- a/test/update-smoke.test.js +++ b/test/update-smoke.test.js @@ -7,6 +7,7 @@ import { tempDir } from "./helpers.js"; import { prepareUpdateSmokeFolder, runUpdateSmoke, + usageLimitedSmoke, validSmokeResponse, } from "../src/gateway/update-smoke.js"; @@ -145,3 +146,38 @@ test("a rejected target is never destroyed and cleanup failures fail verificatio assert.equal(f.calls.at(-1), "release"); rmSync(root, { recursive: true, force: true }); }); + +// Live on Atlas (2026-09-28): the update was refused because the Claude plan had hit its weekly +// limit. A usage-limit answer proves the CLI started in the new container, authenticated and reached +// its provider — the probe's whole purpose — so it counts as reachable, never as a broken gateway. +test("a provider usage limit counts as reachable, thrown or returned; real failures still refuse", async () => { + for (const text of [ + "Claude usage limit reached: You've hit your weekly limit · resets Sep 29, 11pm (Europe/Bucharest)", + "Codex usage limit reached: You've hit your usage limit. Visit chatgpt.com/codex/settings/usage", + "You've hit your session limit · resets 4pm", + "rate_limit_error: too many requests", + ]) assert.equal(usageLimitedSmoke(text), true, text); + for (const text of ["authentication failed", "Engine smoke probe returned an unexpected response.", "ENOENT: claude not found", ""]) { + assert.equal(usageLimitedSmoke(text), false, text); + } + const root = tempRoot(); const f = fixture(root); + try { + const limited = await runUpdateSmoke({ root, resolveTarget: f.resolveTarget, engines: [ + { id: "claude", updateSmoke: async () => { throw new Error("Claude usage limit reached: You've hit your weekly limit · resets Sep 29, 11pm (Europe/Bucharest)"); } }, + { id: "codex", updateSmoke: async () => ({ content: "You've hit your usage limit. Visit chatgpt.com/codex/settings/usage" }) }, + ] }); + assert.equal(limited.ok, true, limited.error); + assert.deepEqual(limited.engines.map((e) => [e.engine, e.ok, e.usageLimited]), [["claude", true, true], ["codex", true, true]]); + assert.match(limited.engines[0].note, /reachable but usage-limited: Claude usage limit reached/); + // A required engine (it passed before the update) that is usage-limited afterwards still passes; + // one that genuinely breaks still refuses. + const after = await runUpdateSmoke({ root, resolveTarget: f.resolveTarget, requiredEngines: ["claude"], engines: [ + { id: "claude", updateSmoke: async () => { throw new Error("Claude usage limit reached: weekly"); } }, + ] }); + assert.equal(after.ok, true); + const broken = await runUpdateSmoke({ root, resolveTarget: f.resolveTarget, engines: [ + { id: "claude", updateSmoke: async () => { throw new Error("authentication failed"); } }, + ] }); + assert.equal(broken.ok, false); + } finally { rmSync(root, { recursive: true, force: true }); } +}); From 0b94153c185bde171050ccbd44f112e4afdf5180 Mon Sep 17 00:00:00 2001 From: Tiberiu Socaci Date: Tue, 29 Sep 2026 10:01:33 +0300 Subject: [PATCH 03/16] fix: contain folder generator test fixtures Signed-off-by: Tiberiu Socaci --- CHANGELOG.md | 1 + TEST-PLAN.md | 9 +++++++++ scripts/run-tests.mjs | 18 +++++++++++------- scripts/test-scratch-sweep.mjs | 19 +++++++++++++------ test/folders-generator-paths.test.js | 19 ++++++++++++++----- test/test-runner-isolation.test.js | 22 ++++++++++++++++++++++ test/test-scratch-cleanup.test.js | 18 ++++++++++++++++++ 7 files changed, 88 insertions(+), 18 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 825cf8ea..1cb1c887 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -18,6 +18,7 @@ product overview. ## Unreleased +- **The folder-generator test no longer scatters projects across the operator's home.** Its custom-workdir fixtures now live under one `ChannelGate Testing` folder, are removed when the test process exits, and stale runs are swept before the next test suite. - **A provider usage limit no longer blocks an update.** Before and after updating, the updater sends each engine a tiny test prompt inside a throwaway container. A reply like "You've hit your weekly limit" used to refuse the update, even though it proves the engine starts, logs in and reaches its provider. It now counts as reachable. A broken login or engine still stops the update. ## 0.6.0 — 2026-09-28 diff --git a/TEST-PLAN.md b/TEST-PLAN.md index b340e5bf..f66e709c 100644 --- a/TEST-PLAN.md +++ b/TEST-PLAN.md @@ -3386,6 +3386,15 @@ exercise the `qwen-eu` adapter itself, in a scratch runtime root (no production hours, keeps a live run's directories and any explicitly protected path, and never touches a directory that is not the suite's. Manual: `ls /tmp | grep -c '^cg-'` before and after a full `npm test` must not grow. +- [x] `test/folders-generator-paths.test.js`: the four custom workdir fixtures are created inside + `~/ChannelGate Testing/folders-generator-*`, never as `cg-*` siblings in the account home. + The process-exit cleanup removes the run folder; `pretest` removes only stale + `folders-generator-*` folders in that parent and leaves recent runs and unrelated folders. + The aggregate runner snapshots that parent before and after the suite and fails on a new + leftover run folder. + Run `node --test test/folders-generator-paths.test.js test/test-scratch-cleanup.test.js`; + pass when both files pass, no `cg-{custom,mirror,project,workdir}-*` directories appear + directly under the home, and no `folders-generator-*` directory remains after the run. - [x] `npm run check:static`: every tracked JavaScript source/test/script parses under the supported Node runtime and fails on tabs or trailing whitespace. This is the deliberately incremental, dependency-free static/format gate; repo-wide ESLint/typed-JS adoption remains a future diff --git a/scripts/run-tests.mjs b/scripts/run-tests.mjs index cc3adba4..0d258ea1 100644 --- a/scripts/run-tests.mjs +++ b/scripts/run-tests.mjs @@ -9,7 +9,7 @@ // Usage: node scripts/run-tests.mjs [--coverage] [extra node flags...] // `--coverage` expands to the coverage flags with the enforced floors below; any other flags are // passed to the spawned `node` in front of `--test`. -import { readdirSync } from "node:fs"; +import { mkdirSync, readdirSync } from "node:fs"; import { fileURLToPath } from "node:url"; import { spawnSync } from "node:child_process"; import os from "node:os"; @@ -53,14 +53,15 @@ const nodeFlags = process.argv .flatMap((flag) => (flag === "--coverage" ? coverageFlags() : [flag])); // ── Real-home leak guard ────────────────────────────────────────────────────────────────────── -// Tests must run entirely in temp directories, pinned through CHANNELGATE_DIR / CHANNELGATE_DB / -// CG_WORKSPACE_DIR (see test/helpers.js). A file that forgets to call ensureTestEnv() silently +// Tests use temp directories, except the one dedicated home folder for custom workdir fixtures. +// Runtime paths are pinned through CHANNELGATE_DIR / CHANNELGATE_DB / CG_WORKSPACE_DIR +// (see test/helpers.js). A file that forgets to call ensureTestEnv() silently // falls back to the DEFAULTS and provisions folders in the operator's real home — which on this // project is a machine running the daemon in production. That is invisible in a green run, so the -// runner snapshots the four root names (current and pre-rename) plus one level beneath each, +// runner snapshots the runtime roots (current and pre-rename) and the fixture parent plus one level beneath each, // before and after, and fails the run on anything new. It is a diff, so directories that leaked in // an earlier run do not fail every run forever — only a NEW entry does. -const HOME_ROOTS = [".channelgate", "ChannelGate", ".claude-gateway", "Slack Agent"]; +const HOME_ROOTS = [".channelgate", "ChannelGate", ".claude-gateway", "Slack Agent", "ChannelGate Testing"]; function homeSnapshot() { const home = os.homedir(); @@ -88,6 +89,9 @@ for (const key of ["CHANNELGATE_DIR", "CHANNELGATE_DB", "CLAUDE_GATEWAY_DIR", "C delete testEnv[key]; } +// The folder-generator test deliberately uses this one named home folder. Include its children +// in the before/after leak guard so a missed cleanup fails the suite. +mkdirSync(path.join(os.homedir(), "ChannelGate Testing"), { recursive: true, mode: 0o700 }); const before = homeSnapshot(); const result = spawnSync(process.execPath, [...nodeFlags, "--test", ...files], { cwd: repoRoot, @@ -103,8 +107,8 @@ if (result.error) { if (leaked.length) { console.error(`\n✖ Tests wrote into the real home directory (${os.homedir()}):`); for (const entry of leaked) console.error(` ${path.join(os.homedir(), entry)}`); - console.error(" Every test must run in a temp directory. Call ensureTestEnv() from test/helpers.js"); - console.error(" before anything resolves a path, or pin CHANNELGATE_DIR / CG_WORKSPACE_DIR yourself."); + console.error(" Use ensureTestEnv() for runtime paths. Register custom workdir fixtures with"); + console.error(" trackTempDir() so ChannelGate Testing is empty after the test process exits."); process.exit(1); } process.exit(result.status ?? 1); diff --git a/scripts/test-scratch-sweep.mjs b/scripts/test-scratch-sweep.mjs index 57487964..ea4785bc 100644 --- a/scripts/test-scratch-sweep.mjs +++ b/scripts/test-scratch-sweep.mjs @@ -23,19 +23,19 @@ import { pathToFileURL } from "node:url"; // The prefixes the suite creates directly under os.tmpdir(): the per-process scratch root, its // TMPDIR sibling, and the workspace roots the container tests pin before ensureTestEnv() runs. -// Everything else a test makes is created AFTER TMPDIR is repointed, so it lives inside a -// `cg-tmp-*` and is removed with it. +// Most other fixtures live inside `cg-tmp-*`; custom workdir fixtures are the one exception, +// grouped under ~/ChannelGate Testing and swept separately below. export const SCRATCH_PREFIXES = Object.freeze(["cg-test-", "cg-tmp-", "cg-ws-"]); // Old enough that no live run can own it. Directory mtime moves whenever a direct child is added // or removed, so an active scratch root keeps refreshing itself well inside this window. export const MAX_AGE_MS = 2 * 60 * 60 * 1000; -export function isScratchName(name) { - return SCRATCH_PREFIXES.some((prefix) => name.startsWith(prefix)); +export function isScratchName(name, prefixes = SCRATCH_PREFIXES) { + return prefixes.some((prefix) => name.startsWith(prefix)); } -export function sweepScratchDirs({ root = os.tmpdir(), now = Date.now(), maxAgeMs = MAX_AGE_MS, keep = [] } = {}) { +export function sweepScratchDirs({ root = os.tmpdir(), now = Date.now(), maxAgeMs = MAX_AGE_MS, keep = [], prefixes = SCRATCH_PREFIXES } = {}) { const result = { root, removed: 0, recent: 0, skipped: 0 }; // Never touch a directory this very process was pointed at. const protectedDirs = new Set(keep.filter(Boolean).map((dir) => path.resolve(dir))); @@ -47,7 +47,7 @@ export function sweepScratchDirs({ root = os.tmpdir(), now = Date.now(), maxAgeM } const cutoff = now - maxAgeMs; for (const entry of entries) { - if (!entry.isDirectory() || !isScratchName(entry.name)) continue; + if (!entry.isDirectory() || !isScratchName(entry.name, prefixes)) continue; const abs = path.join(root, entry.name); if (protectedDirs.has(path.resolve(abs))) continue; const stats = statSync(abs, { throwIfNoEntry: false }); @@ -68,10 +68,17 @@ export function sweepScratchDirs({ root = os.tmpdir(), now = Date.now(), maxAgeM function main() { const swept = sweepScratchDirs({ keep: [process.env.TMPDIR, process.env.CG_TEST_SCRATCH] }); + const fixtures = sweepScratchDirs({ + root: path.join(os.homedir(), "ChannelGate Testing"), + prefixes: ["folders-generator-"], + }); const parts = [`removed ${swept.removed}`]; if (swept.recent) parts.push(`kept ${swept.recent} recent`); if (swept.skipped) parts.push(`skipped ${swept.skipped} not removable`); console.log(`[test-scratch-sweep] ${parts.join(", ")} in ${swept.root}`); + if (fixtures.removed || fixtures.recent || fixtures.skipped) { + console.log(`[test-scratch-sweep] removed ${fixtures.removed}, kept ${fixtures.recent} recent, skipped ${fixtures.skipped} in ${fixtures.root}`); + } } const invoked = process.argv[1] && import.meta.url === pathToFileURL(path.resolve(process.argv[1])).href; diff --git a/test/folders-generator-paths.test.js b/test/folders-generator-paths.test.js index c6331519..3f00e93a 100644 --- a/test/folders-generator-paths.test.js +++ b/test/folders-generator-paths.test.js @@ -8,9 +8,12 @@ import os from "node:os"; import path from "node:path"; import { test } from "node:test"; import assert from "node:assert/strict"; -import { ensureTestEnv } from "./helpers.js"; +import { ensureTestEnv, trackTempDir } from "./helpers.js"; ensureTestEnv(); +// This test checks custom paths under the account's allowed root, regardless of a developer's +// CG_FS_ROOT override. Its disposable projects are grouped below that home instead of beside it. +process.env.CG_FS_ROOT = os.homedir(); const [ { effectiveWorkDir, updateChannelInstructions, listAvailableSkills, skillSourceDirs, enableSkills, splitGatewayBlock, gatewayInstructionsBlock, channelSwitchesNote }, @@ -22,12 +25,18 @@ const [ import("../src/web/security.js"), ]); +// Custom workDirs must live under the allowed filesystem root. Keep this test's disposable +// projects together instead of scattering cg-* directories across the operator's home. +const fixtureParent = path.join(allowedFsRoot(), "ChannelGate Testing"); +mkdirSync(fixtureParent, { recursive: true, mode: 0o700 }); +const fixtureRoot = trackTempDir(mkdtempSync(path.join(fixtureParent, "folders-generator-"))); + test("effectiveWorkDir: clean mode, a contained custom workDir, and an escaping one", () => { const slug = "workdir-probe"; assert.equal(effectiveWorkDir(slug, { cleanMode: true }), cleanWorkspaceFolder(slug, undefined)); // A stored absolute workDir inside the allowed root that exists is honoured as-is. - const inside = mkdtempSync(path.join(allowedFsRoot(), "cg-workdir-")); + const inside = mkdtempSync(path.join(fixtureRoot, "cg-workdir-")); assert.equal(effectiveWorkDir(slug, { workDir: inside }), inside); // The same path once it has vanished falls back to the default folder. @@ -68,7 +77,7 @@ test("updateChannelInstructions keeps the managed block on a default folder and assert.match(content, /Only this rule\.\n$/); // Custom project folder: the file's own content is edited, no block is added. - const custom = mkdtempSync(path.join(allowedFsRoot(), "cg-custom-")); + const custom = mkdtempSync(path.join(fixtureRoot, "cg-custom-")); const customMeta = { ...meta, workDir: custom }; await updateChannelInstructions(slug, customMeta, { text: "Project rule one." }); await updateChannelInstructions(slug, customMeta, { text: "Project rule two." }); @@ -78,7 +87,7 @@ test("updateChannelInstructions keeps the managed block on a default folder and // A project that keeps AGENTS.md as the real file with the gateway's mirror link on CLAUDE.md // is edited through the link's exact sibling shape only. - const mirrored = mkdtempSync(path.join(allowedFsRoot(), "cg-mirror-")); + const mirrored = mkdtempSync(path.join(fixtureRoot, "cg-mirror-")); writeFileSync(path.join(mirrored, "AGENTS.md"), "# Project\n"); const { symlinkSync } = await import("node:fs"); symlinkSync("AGENTS.md", path.join(mirrored, "CLAUDE.md")); @@ -127,7 +136,7 @@ test("skill listing unions the catalog with the host folders, and a host-folder test("a custom project folder gets a real CLAUDE.md with AGENTS.md mirrored, and a renamed managed skill folder is migrated", async () => { const { ensureChannelFolder } = await import("../src/gateway/folders.js"); const { MANAGED_SKILL_MARKER } = await import("../src/gateway/skills/materialize.js"); - const custom = mkdtempSync(path.join(allowedFsRoot(), "cg-project-")); + const custom = mkdtempSync(path.join(fixtureRoot, "cg-project-")); const slug = "custom-project-probe"; await ensureChannelFolder(slug, { name: "Custom project", workDir: custom, allowedMcps: [], instructions: "Project instructions from the channel record." }); const claude = path.join(custom, "CLAUDE.md"); diff --git a/test/test-runner-isolation.test.js b/test/test-runner-isolation.test.js index 198db306..83ac9394 100644 --- a/test/test-runner-isolation.test.js +++ b/test/test-runner-isolation.test.js @@ -96,3 +96,25 @@ test("aggregate test runner removes serving-runtime selectors while preserving o assert.ok(observed.flags.includes("--no-warnings")); assert.deepEqual(snapshot(external), before); }); + +test("aggregate test runner detects a fixture left in ChannelGate Testing", () => { + const fixture = tempDir("cg-suite-fixture-leak-"); + mkdirSync(path.join(fixture, "scripts")); + mkdirSync(path.join(fixture, "test")); + copyFileSync(path.join(repoRoot, "scripts", "run-tests.mjs"), path.join(fixture, "scripts", "run-tests.mjs")); + writeFileSync(path.join(fixture, "package.json"), '{"type":"module"}'); + writeFileSync(path.join(fixture, "test", "leak.test.js"), ` + import { mkdirSync } from 'node:fs'; + import os from 'node:os'; + import path from 'node:path'; + mkdirSync(path.join(os.homedir(), 'ChannelGate Testing', 'folders-generator-leak')); + `); + const env = { ...process.env, HOME: fixture }; + delete env.NODE_TEST_CONTEXT; + const result = spawnSync(process.execPath, ["scripts/run-tests.mjs"], { + cwd: fixture, env, encoding: "utf8", timeout: 60_000, + }); + assert.equal(result.status, 1); + assert.match(result.stderr, /Tests wrote into the real home directory/); + assert.match(result.stderr, /ChannelGate Testing\/folders-generator-leak/); +}); diff --git a/test/test-scratch-cleanup.test.js b/test/test-scratch-cleanup.test.js index 5cb810be..b73126c7 100644 --- a/test/test-scratch-cleanup.test.js +++ b/test/test-scratch-cleanup.test.js @@ -85,3 +85,21 @@ test("the sweeper removes stale scratch roots, keeps live ones, and never touche assert.equal(sweepScratchDirs({ root, keep: [liveRoot] }).removed, 0); assert.equal(existsSync(liveRoot), true); }); + +test("the dedicated fixture folder only sweeps stale generator runs", () => { + const root = tempDir("cg-fixture-sweep-"); + const old = Date.now() / 1000 - 6 * 60 * 60; + const stale = path.join(root, "folders-generator-old"); + const recent = path.join(root, "folders-generator-live"); + const other = path.join(root, "keep-this-project"); + for (const dir of [stale, recent, other]) mkdirSync(dir); + utimesSync(stale, old, old); + utimesSync(other, old, old); + + const swept = sweepScratchDirs({ root, prefixes: ["folders-generator-"] }); + assert.equal(swept.removed, 1); + assert.equal(swept.recent, 1); + assert.equal(existsSync(stale), false); + assert.equal(existsSync(recent), true); + assert.equal(existsSync(other), true); +}); From e8641df640f5f9ff4446e5d3d8a4cc24eb852877 Mon Sep 17 00:00:00 2001 From: Tiberiu Socaci Date: Tue, 29 Sep 2026 10:50:39 +0300 Subject: [PATCH 04/16] feat(admin): finish Overview model charts and layout Signed-off-by: Tiberiu Socaci --- CHANGELOG.md | 4 ++ FEATURES.md | 11 +++-- TEST-PLAN.md | 16 +++++- public/app.js | 76 ++++++++++++++--------------- public/styles.css | 22 +++++++-- test/dashboard-chart-layout.test.js | 65 ++++++++++++++++++++++++ test/usage-model-breakdown.test.js | 2 +- 7 files changed, 148 insertions(+), 48 deletions(-) create mode 100644 test/dashboard-chart-layout.test.js diff --git a/CHANGELOG.md b/CHANGELOG.md index 1cb1c887..12143388 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -18,6 +18,10 @@ product overview. ## Unreleased +- Overview now draws separate stacked columns for each time bucket, shows model totals when a + source, channel or user bar is hovered or focused, and arranges usage sources beside the three + time charts above Models, Channels, Users and Skills. + - **The folder-generator test no longer scatters projects across the operator's home.** Its custom-workdir fixtures now live under one `ChannelGate Testing` folder, are removed when the test process exits, and stale runs are swept before the next test suite. - **A provider usage limit no longer blocks an update.** Before and after updating, the updater sends each engine a tiny test prompt inside a throwaway container. A reply like "You've hit your weekly limit" used to refuse the update, even though it proves the engine starts, logs in and reaches its provider. It now counts as reachable. A broken login or engine still stops the update. diff --git a/FEATURES.md b/FEATURES.md index e9c72885..e5d655f0 100644 --- a/FEATURES.md +++ b/FEATURES.md @@ -3832,13 +3832,13 @@ are retired, bullet by bullet; everything else stands. active users, active sessions, and total tokens), per-bucket charts for sessions/tokens/cost, a sessions-per-user bar list (descending), and a per-channel sessions+cost bar chart. Pure inline SVG + div bars — no chart library, no build step. → TEST-PLAN: Admin UI. -- **Every Overview chart is stacked by the model that answered.** A cost, run or token line split +- **Every Overview time chart is stacked by the model that answered.** A cost, run or token column split by model answers "did spend rise because we ran more, or because we moved onto a pricier model" — which one undifferentiated line cannot. Attribution is per component, not per run: a Codex turn whose subagent ran a different model contributes a slice to EACH, while the run itself is still counted once (`MODEL_ATTRIBUTION_CTE` in `src/gateway/usage.js`). The split reconciles with the headline by construction — a run whose components are only partly priced contributes no dollars - to any model, exactly as the canonical rollup drops it — so the bands always add up to the total + to any model, exactly as the canonical rollup drops it — so the columns always add up to the total beside them. Model ids are normalised for display (`modelDisplayLabel`): a dated snapshot (`claude-haiku-4-5-20251001`) reads as its family, a context variant keeps its `1M` marker, an OpenAI id becomes `GPT-5.6 Sol`, and an unrecognised id passes through verbatim rather than being @@ -3858,13 +3858,18 @@ are retired, bullet by bullet; everything else stands. - **A dedicated Models chart**, answering "which models are used more" directly: every model in the window as a bar — token cost, its share of spend, runs (and how many of them came from outside the gateway) and tokens. The stacked charts above cap at the seven largest models plus a neutral - "Other" band, because a stacked area with a generated ninth hue stops being readable; this card + "Other" segment, because a generated ninth hue stops being readable; this card lists everything, so nothing is hidden by that cap. Colours come from a categorical palette validated for the admin surface (lightness band, chroma floor, adjacent colourblind separation, normal-vision separation and 3:1 contrast all pass) and are keyed on the MODEL, so changing range, harness or source never repaints the series that survived the filter. Each chart carries a legend and a crosshair tooltip breaking the hovered bucket down per model (pointer or keyboard). → TEST-PLAN: Admin UI. +- **Overview chart order and details:** four compact cards show token cost, runs, tokens and usage + sources. Below them, full-width cards show Models, Channels, Users and Skills in that order. + Time buckets draw stacked columns. Hovering or keyboard-focusing a source, channel or user bar + shows the total and its per-model values for that bar's metric; the same model colours and labels + appear in the time-chart tooltips and legend. → TEST-PLAN: Per-model dashboard stacking. - **The bar lists stack by model too**, in the same colours and the same order as the charts, so one hue means one model across the whole page: Runs per user (stacked by runs), Channels (all three bars — runs, token cost, tokens — each stacked by its own metric), and Where usage came from diff --git a/TEST-PLAN.md b/TEST-PLAN.md index f66e709c..45b8d2b0 100644 --- a/TEST-PLAN.md +++ b/TEST-PLAN.md @@ -1221,9 +1221,23 @@ unchecked live gate above. snapshot, OpenAI id, unknown id, empty). Engine-independent: this is ledger SQL, not harness behaviour. - [x] `test/usage-model-breakdown.test.js` also pins the Overview's chart contract: the categorical - palette hexes (changing one obliges re-running the dataviz validator), the stacked-area + palette hexes (changing one obliges re-running the dataviz validator), the stacked-column renderer, a legend for every multi-series chart, the models bar chart, and that hues are keyed on the model rather than cycled by position in a filtered list. +- [x] `test/dashboard-chart-layout.test.js`: stacked columns occupy separate time buckets and + reconcile their heights with per-model values; source, channel and user bars expose the + matching model breakdown on hover and keyboard focus; the four compact cards and four + full-width cards render in the requested order. Engine-independent: these are browser UI + functions over the already-aggregated dashboard payload. +- [ ] Live acceptance (Overview chart layout): open Overview with a range containing usage from + two or more models. Require four cards across at desktop width: Token est. cost, Runs, + Tokens, Where usage came from. Require separate stacked columns for each time bucket and + verify the hovered column shows that bucket's model values. Below, require full-width Models, + Channels, Runs per user, Top skills in that order. Hover and keyboard-focus one source bar, + one channel metric bar and one user bar; each tooltip must show the bar's total and model + breakdown in the same colours as the legend. Repeat at a narrow viewport to check card + wrapping and that tooltips remain readable. Engine-independent: the UI reads a fixed API + payload and no engine turn is involved. - [x] `test/external-usage.test.js`: Claude transcript parsing into per-hour/per-model aggregates (subagent spend counted, synthetic error replies not, tool results and subagent prompts not counted as turns, dated snapshots collapsed onto the billing id); the cache-write TTL split, diff --git a/public/app.js b/public/app.js index f9c92e79..21c97c54 100644 --- a/public/app.js +++ b/public/app.js @@ -683,38 +683,34 @@ function stackValue(point, entry, metric) { return entry.rest.reduce((total, model) => total + (Number(models[model]?.[metric]) || 0), 0); } -// Stacked area chart. Same stretched viewBox and non-scaling strokes as sparkArea, so it drops into -// the existing chart cards unchanged. Bands are separated by a 2px stroke in the CARD's own colour -// rather than a gap in the geometry: at one-pixel bucket widths a geometric gap would swallow thin -// series whole. -function stackedArea(series, keys, metric, opts = {}) { +// One stacked column per time bucket. Keep the columns centered within their buckets so the hover +// target and the visible bar always refer to the same day, hour or month. +function stackedColumns(series, keys, metric, opts = {}) { const W = 300, H = opts.height || 110, pad = 4; const n = series.length; const cls = "spark" + (opts.tall ? " spark-tall" : ""); if (!n || !keys.length) return ``; const totals = series.map((point) => keys.reduce((sum, entry) => sum + stackValue(point, entry, metric), 0)); const max = Math.max(1e-9, ...totals); - const xs = (i) => (n <= 1 ? W / 2 : (i / (n - 1)) * W); const ys = (v) => H - pad - (v / max) * (H - pad * 2); + const step = W / n; + const width = Math.max(1, step - Math.min(2, step * 0.2)); const grid = opts.grid ? [1, 2].map((k) => ``).join("") : ""; - // Cumulative from the baseline up, so each band's lower edge is the previous band's upper edge. const running = new Array(n).fill(0); - const bands = []; + const columns = []; for (const entry of keys) { - const lower = running.map((v) => v); - for (let i = 0; i < n; i++) running[i] += stackValue(series[i], entry, metric); - const upper = running.map((v) => v); - if (upper.every((v, i) => v === lower[i])) continue; // a model with nothing in this window - const top = upper.map((v, i) => `${xs(i).toFixed(1)},${ys(v).toFixed(1)}`); - const bottom = lower.map((v, i) => `${xs(i).toFixed(1)},${ys(v).toFixed(1)}`).reverse(); - bands.push( - `` + - `` - ); + for (let i = 0; i < n; i++) { + const value = stackValue(series[i], entry, metric); + if (value <= 0) continue; + const bottom = ys(running[i]); + running[i] += value; + const top = ys(running[i]); + columns.push(``); + } } - return `${grid}${bands.join("")}`; + return `${grid}${columns.join("")}`; } // Legend for a stacked chart. Always rendered when there is more than one series — identity must @@ -726,13 +722,12 @@ function modelLegend(keys) { .join("")}`; } -// A stacked chart card, with the hover layer the plain sparkline cards do not need: an area chart -// that stacks eight series is unreadable without being able to ask "what is this band, here". +// A stacked chart card with a hover layer for the model totals in each time bucket. function stackedChartCard(id, title, series, keys, metric, peak, axis, fmt, opts = {}) { return `

${escapeHtml(title)}

${peak}
- ${stackedArea(series, keys, metric, opts)} + ${stackedColumns(series, keys, metric, opts)}
@@ -754,7 +749,14 @@ function barTrack(pct, { color = "", stack = null } = {}) { const segments = parts .map((part) => ``) .join(""); - return `${segments}`; + const format = stack.metric === "cost" ? fmtUSD : stack.metric === "tokens" ? fmtCompact : fmtNum; + const metricName = { cost: "Token cost", tokens: "Tokens", runs: "Runs" }[stack.metric] || stack.metric; + const details = parts.map((part) => `${escapeHtml(part.entry.label)}${format(part.value)}`).join(""); + const accessible = `${stack.row.name || "Usage"}, ${metricName}: ${format(total)}. ${parts.map((part) => `${part.entry.label}: ${format(part.value)}`).join(", ")}`; + return ` + ${segments} + ${escapeHtml(metricName)} · ${format(total)}${details} + `; } // Single-metric horizontal bar list, descending. Row: name · bar (width ∝ value) · value. @@ -1100,13 +1102,12 @@ async function loadDashboard() {
`; }).join(""); - // Token cost is the hero (2fr wide, gridlines, peak dated). Runs + tokens ride at 1fr but share the - // hero's chart height so all three axis labels line up along the same bottom edge. + // Four compact cards share one row: cost, runs, tokens and the usage-source breakdown. const costPeak = peakBucket((x) => x.cost); const costPeakLabel = costPeak && costPeak.cost > 0 ? `peak ${fmtUSD(costPeak.cost)}${costPeak.key ? " · " + bucketLabel(costPeak.key, unit) : ""}` : "no value yet"; - // Every hero chart is stacked by the model that actually answered, so a rising cost line can be + // Every time chart is stacked by the model that actually answered, so a rising cost column can be // read as "we moved onto a pricier model" rather than only "we ran more". One colour map and one // key list across all three, so a band means the same thing in each and the legend is shared. const allModels = d.models || []; @@ -1164,15 +1165,9 @@ async function loadDashboard() {
${kpiHtml}
${modelLegend(keys)}
-
${charts}
-
+
${charts}${originsCard}
+
${modelsCard} - ${originsCard} -
-

Runs per user

${users.length} of ${fmtNum(t.users)}
- ${barList(users, (u) => u.runs, (u) => `${fmtNum(u.runs)} runs${fmtCompact(u.tokens)} tokens${fmtUSD(u.cost)} est.`, "#91c9ce", "No user activity yet.", { keys, metric: "runs" })} - ${moreUsers > 0 ? `

+ ${moreUsers} more

` : ""} -

Channels — runs, token cost & tokens

three bars per channel, split by model @@ -1180,6 +1175,11 @@ async function loadDashboard() { ${channelBars(channels, keys)} ${moreChannels > 0 ? `

+ ${moreChannels} more

` : ""}
+
+

Runs per user

${users.length} of ${fmtNum(t.users)}
+ ${barList(users, (u) => u.runs, (u) => `${fmtNum(u.runs)} runs${fmtCompact(u.tokens)} tokens${fmtUSD(u.cost)} est.`, "#91c9ce", "No user activity yet.", { keys, metric: "runs" })} + ${moreUsers > 0 ? `

+ ${moreUsers} more

` : ""} +

Top skills — usage

last 30 days
${barList(skills, (s) => s.uses, (s) => `${fmtNum(s.uses)} uses`, "#6ea6a1", "No skill usage yet.")} @@ -1209,8 +1209,8 @@ async function loadDashboard() { } } -// Crosshair + tooltip for the stacked charts. A stacked area with up to eight bands cannot be read -// without asking "which band is this, and how much"; the legend names the colours, this says the +// Crosshair + tooltip for the stacked time charts. A column with up to eight segments cannot be read +// without asking "which model is this, and how much"; the legend names the colours, this says the // numbers. Pointer-driven and keyboard-reachable (the chart is focusable and arrow keys step // buckets), so the reading is not mouse-only. const STACK_FMT = { usd: (v) => fmtUSD(v), num: (v) => fmtNum(v), compact: (v) => fmtCompact(v) }; @@ -1228,7 +1228,7 @@ function renderStackTip(host, index) { const total = rows.reduce((sum, row) => sum + row.value, 0); const tip = host.querySelector(".spark-tip"); const cursor = host.querySelector(".spark-cursor"); - const pct = data.series.length <= 1 ? 50 : (index / (data.series.length - 1)) * 100; + const pct = ((index + 0.5) / data.series.length) * 100; cursor.style.left = `${pct}%`; cursor.hidden = false; tip.hidden = false; @@ -1262,7 +1262,7 @@ function wireStackHover(root) { if (!n) return; const rect = host.getBoundingClientRect(); if (!rect.width) return; - show(Math.round(((event.clientX - rect.left) / rect.width) * (n - 1))); + show(Math.floor(((event.clientX - rect.left) / rect.width) * n)); }); host.addEventListener("pointerleave", hide); host.addEventListener("focus", () => show(current >= 0 ? current : count() - 1)); diff --git a/public/styles.css b/public/styles.css index 5d76e7cf..183f6fae 100644 --- a/public/styles.css +++ b/public/styles.css @@ -412,9 +412,9 @@ button.clear-tok.armed { background: rgba(229, 96, 77, .14); border-color: rgba( .kpi-link:hover { border-color: var(--orange); background: var(--panel-2); transform: translateY(-1px); } .kpi-link:focus-visible { outline: none; border-color: var(--orange); } -/* Cost is the hero chart (2fr); sessions + tokens ride at 1fr each. */ -.dash-grid { display: grid; grid-template-columns: 2fr 1fr 1fr; gap: 16px; margin-bottom: 20px; } -.dash-two { display: grid; grid-template-columns: 1fr 1fr; gap: 16px; } +/* Four compact usage charts above the full-width model, channel, user and skill cards. */ +.dash-grid { display: grid; grid-template-columns: repeat(4, minmax(0, 1fr)); gap: 16px; margin-bottom: 20px; } +.dash-stack { display: grid; grid-template-columns: minmax(0, 1fr); gap: 16px; } .chart-card { background: var(--panel); border: 1px solid var(--line-soft); border-radius: var(--radius); padding: 16px 18px; min-width: 0; } .chart-title { display: flex; align-items: baseline; justify-content: space-between; gap: 8px; margin-bottom: 10px; } @@ -427,7 +427,7 @@ button.clear-tok.armed { background: rgba(229, 96, 77, .14); border-color: rgba( .legend .dot { display: inline-block; width: 9px; height: 9px; border-radius: 2px; margin-right: 5px; vertical-align: middle; box-shadow: none; } /* ── Per-model stacking ────────────────────────────────────────────────────── - One legend above the three hero charts: the bands mean the same model in each, + One legend above the four usage cards: the colours mean the same model in each, so repeating it per card would be three copies of the same key. It wraps because a deployment can easily be running eight models at once. Legend text stays in the muted ink token — the swatch carries identity, never the label's colour. */ @@ -457,6 +457,13 @@ button.clear-tok.armed { background: rgba(229, 96, 77, .14); border-color: rgba( .spark-tip-row b { margin-left: auto; font-variant-numeric: tabular-nums; } .spark-tip-row .dot { display: inline-block; width: 8px; height: 8px; border-radius: 2px; flex: none; } +/* Model breakdown for each source/user bar and each of a channel's three metrics. */ +.bar-hover { position: relative; display: block; min-width: 0; cursor: help; } +.bar-hover:focus-visible { outline: 2px solid var(--orange); outline-offset: 2px; border-radius: 5px; } +.bar-tip { display: none; right: 0; bottom: calc(100% + 6px); max-width: min(260px, 85vw); } +.bar-hover:hover .bar-tip, .bar-hover:focus .bar-tip { display: block; } +.bar-tip .spark-tip-head { display: block; } + .bar-row { display: grid; grid-template-columns: 120px 1fr auto; align-items: center; gap: 10px; padding: 6px 0; font-size: 12.5px; } .bar-row-dual { align-items: center; } .bar-name { overflow: hidden; text-overflow: ellipsis; white-space: nowrap; color: var(--ink); } @@ -476,6 +483,8 @@ button.clear-tok.armed { background: rgba(229, 96, 77, .14); border-color: rgba( /* Stacked value: when a bar's value is split into lines (e.g. sessions/tokens/cost per user), render one metric per line. Plain-text values (channel rows) are unaffected. */ .bar-val > span { display: block; line-height: 1.45; } +.dash-grid .bar-row { grid-template-columns: 76px minmax(24px, 1fr) auto; gap: 6px; font-size: 11px; } +.dash-grid .bar-val { font-size: 10px; } /* ── Overview (slice 2): orange cost KPI + taller hero cost chart ───────────── */ .kpi-value.cost { color: var(--orange); } @@ -869,7 +878,7 @@ button.clear-tok.armed { background: rgba(229, 96, 77, .14); border-color: rgba( /* ── Responsive: turn the left rail into a compact row above the sections ─── */ @media (max-width: 900px) { - .dash-grid, .dash-two { grid-template-columns: 1fr; } + .dash-grid { grid-template-columns: repeat(2, minmax(0, 1fr)); } .cap-cards { grid-template-columns: repeat(2, 1fr); } .userlayout { flex-direction: column; } .userlayout .drawer { width: 100%; position: static; } @@ -880,6 +889,9 @@ button.clear-tok.armed { background: rgba(229, 96, 77, .14); border-color: rgba( .setcontent { width: 100%; } .setbody { max-width: none; } } +@media (max-width: 600px) { + .dash-grid { grid-template-columns: minmax(0, 1fr); } +} @media (max-width: 760px) { .app { flex-direction: column; height: auto; } diff --git a/test/dashboard-chart-layout.test.js b/test/dashboard-chart-layout.test.js new file mode 100644 index 00000000..73b8af87 --- /dev/null +++ b/test/dashboard-chart-layout.test.js @@ -0,0 +1,65 @@ +import test from "node:test"; +import assert from "node:assert/strict"; +import { readFile } from "node:fs/promises"; +import { runInNewContext } from "node:vm"; + +const app = await readFile(new URL("../public/app.js", import.meta.url), "utf8"); +const css = await readFile(new URL("../public/styles.css", import.meta.url), "utf8"); +const between = (start, end) => app.slice(app.indexOf(start), app.indexOf(end)); + +test("time buckets render separate stacked columns with model segments", () => { + const chart = runInNewContext(`${between("function stackValue(", "// Legend for a stacked chart")}; stackedColumns`, {}); + const keys = [ + { key: "a", color: "#111111" }, + { key: "b", color: "#222222" }, + ]; + const series = [ + { models: { a: { cost: 3 }, b: { cost: 2 } } }, + { models: { a: { cost: 1 }, b: { cost: 4 } } }, + ]; + const svg = chart(series, keys, "cost", { height: 110 }); + const rects = [...svg.matchAll(//g)] + .map((match) => ({ x: Number(match[1]), y: Number(match[2]), width: Number(match[3]), height: Number(match[4]), color: match[5] })); + assert.equal(rects.length, 4); + assert.deepEqual([...new Set(rects.map((r) => r.x))], [1, 151]); + assert.ok(rects.every((r) => r.width < 150 && r.height > 0)); + for (const x of [1, 151]) { + const column = rects.filter((r) => r.x === x); + assert.deepEqual(column.map((r) => r.color), ["#111111", "#222222"]); + assert.ok(Math.abs(column.reduce((sum, r) => sum + r.height, 0) - 102) < 0.02); + } +}); + +test("model bars expose the total and breakdown on hover and keyboard focus", () => { + const render = runInNewContext(`${between("function stackValue(", "// Legend for a stacked chart")} + ${between("function barTrack(", "// Single-metric horizontal bar list")}; barTrack`, { + fmtUSD: (n) => `$${n.toFixed(2)}`, + fmtCompact: String, + fmtNum: String, + escapeHtml: (s) => String(s).replaceAll("&", "&").replaceAll('"', """), + MODEL_OTHER: "#888888", + }); + const html = render(80, { + stack: { + row: { name: "Test channel", models: { a: { runs: 3 }, b: { runs: 2 } } }, + keys: [{ key: "a", label: "Model A", color: "#111111" }, { key: "b", label: "Model B", color: "#222222" }], + metric: "runs", + }, + }); + assert.match(html, /tabindex="0"/); + assert.match(html, /Test channel, Runs: 5\. Model A: 3, Model B: 2/); + assert.match(html, /class="spark-tip bar-tip"/); + assert.match(html, /Model A3<\/b>/); + assert.match(html, /Model B2<\/b>/); + assert.match(css, /\.bar-hover:hover \.bar-tip, \.bar-hover:focus \.bar-tip \{ display: block; \}/); +}); + +test("Overview orders four compact cards above Models, Channels, Users and Skills", () => { + assert.match(app, /
\$\{charts\}\$\{originsCard\}<\/div>/); + const stack = between('
', "// Approvals waiting on a human"); + const order = ["${modelsCard}", "Channels —", "Runs per user", "Top skills —"].map((label) => stack.indexOf(label)); + assert.ok(order.every((index) => index >= 0)); + assert.deepEqual([...order].sort((a, b) => a - b), order); + assert.match(css, /\.dash-grid \{[^}]*repeat\(4, minmax\(0, 1fr\)\)/); + assert.match(css, /\.dash-stack \{[^}]*grid-template-columns: minmax\(0, 1fr\)/); +}); diff --git a/test/usage-model-breakdown.test.js b/test/usage-model-breakdown.test.js index ba2f1757..c68d8d00 100644 --- a/test/usage-model-breakdown.test.js +++ b/test/usage-model-breakdown.test.js @@ -158,7 +158,7 @@ test("the Overview ships a stacked-by-model chart, a legend and a dedicated mode // The validated categorical palette (scripts/validate_palette.js from the dataviz skill). If a // hue changes here, that validator has to be re-run — this assertion is the reminder. assert.match(app, /const MODEL_PALETTE = \["#3987e5", "#d95926", "#199e70", "#c98500", "#d55181", "#48a02b", "#9085e9", "#e66767"\];/); - assert.match(app, /function stackedArea\(/); + assert.match(app, /function stackedColumns\(/); assert.match(app, /function modelLegend\(/, "identity is never carried by colour alone"); assert.match(app, /function modelBars\(/); // The bar lists stack by the same model keys and colours the charts use, so one hue means one From 05ad9277307cf407a7a49e6477c489e093420793 Mon Sep 17 00:00:00 2001 From: Tiberiu Socaci Date: Tue, 29 Sep 2026 14:29:35 +0300 Subject: [PATCH 05/16] fix: keep successful engine failovers on their threads Signed-off-by: Tiberiu Socaci --- CHANGELOG.md | 5 +++ FEATURES.md | 6 ++-- TEST-PLAN.md | 13 +++++++ src/gateway/run.js | 30 ++++++++++++---- src/gateway/sessions.js | 37 ++++++++++++++++--- test/claude-fallback-e2e.test.js | 49 ++++++++++++++++++++++++++ test/codex-failover-e2e.test.js | 11 +++++- test/runtime-identity-preamble.test.js | 4 +-- 8 files changed, 139 insertions(+), 16 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 12143388..6a1d03b7 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -18,6 +18,11 @@ product overview. ## Unreleased +- **Automatic engine switches now stay with the conversation.** When Claude reaches its limit and + Codex answers, the next message resumes that Codex session instead of retrying Claude. The same + holds when Codex switches to Claude. Existing threads with a successful fallback session adopt + it on their next message; a manual thread engine or model choice still takes precedence. + - Overview now draws separate stacked columns for each time bucket, shows model totals when a source, channel or user bar is hovered or focused, and arranges usage sources beside the three time charts above Models, Channels, Users and Skills. diff --git a/FEATURES.md b/FEATURES.md index e5d655f0..9f2a9d23 100644 --- a/FEATURES.md +++ b/FEATURES.md @@ -1451,8 +1451,10 @@ A categorized catalog of what's shipped. Cross-linked to `TEST-PLAN.md` checks. (OpenAI Codex CLI — one-shot per message, MCP via `-c` overrides + an HTTP bridge). Selection precedence is **per-thread directive → per-channel → global default** — but an EXISTING thread sticks to the engine that minted its session: changing the channel/global harness only affects - new threads; live conversations keep resuming on their own engine (only the automatic usage-limit - failover runs a different engine, under a suffixed session key). The exception is a harness turned + new threads; live conversations keep resuming on their own engine. A successful automatic + cross-engine failover makes the answering engine the thread's live session, so its next turn + resumes there even after the failed engine's cooldown ends. Older fallback-only sessions are + adopted on their next unpinned turn. The exception is a harness turned OFF in Settings: those threads move to an enabled engine. A row minted for a brand-new thread whose turn then dies BEFORE its engine starts (a pre-spawn credential gate, a runtime that cannot come up) is dropped with that turn, so the next message is a first turn again on the channel's diff --git a/TEST-PLAN.md b/TEST-PLAN.md index 45b8d2b0..a3af8f81 100644 --- a/TEST-PLAN.md +++ b/TEST-PLAN.md @@ -4509,6 +4509,19 @@ placeholders and `--network none`). authentication failure answers via the OTHER harness with a reason note and observes its per-engine per-channel / gateway-wide ~15-min cooldown; with failover OFF, the engine's own error surfaces. A post-tool failure never replays. +- [x] Automatic failover stays on the answering harness (`test/claude-fallback-e2e.test.js`, + `test/codex-failover-e2e.test.js`): in a Claude-default channel, trigger the fixture's + replay-safe Claude limit and let Codex answer; reset cooldown and send another message in + the same thread. Pass: the second turn runs on Codex with `resume=yes`, with no new failover + note. Repeat in a Codex-default channel with the Codex limit and Claude answer. For an older + session fixture with a Claude main row and a successful Codex fallback row, send an unpinned + continuation; pass: it resumes Codex and moves that session to the main thread key. A manual + thread engine/model choice keeps its existing precedence. Live acceptance on either engine: + use a test account whose primary harness is genuinely limited, observe the fallback answer, + then send a second message in the same thread after its cooldown; pass only if the second + reply footer names the fallback harness and continues its prior context. Private Airtable + live definitions `ENG-11` (Claude→Codex on Atlas) and `ENG-12` (Codex→Claude on Xavier) + are registered and remain unexecuted until their limited-account fixtures are available. - [x] Unit: the Codex runner classifies its plan-limit rejection ("purchase more credits…") as a replay-safe `usage_limit` — as a JSON error event AND on stderr with a nonzero exit — while model rejections keep routing to the same-engine model retry, server/connection errors diff --git a/src/gateway/run.js b/src/gateway/run.js index 4f593c95..08cfe209 100644 --- a/src/gateway/run.js +++ b/src/gateway/run.js @@ -14,7 +14,7 @@ import { memorySnapshotPrefix } from "./channel-memory.js"; import { recordUsage } from "./usage.js"; import { createSkillUsageRecorder } from "./skills/usage.js"; import { withTemplateSkills } from "./skills/templates.js"; -import { resolveSession, resetSession, getSession, saveSession, sessionGeneration, dropMintedSession } from "./sessions.js"; +import { resolveSession, resetSession, getSession, getSessionEngine, saveSession, promoteFallbackSession, discardFallbackSessions, sessionGeneration, dropMintedSession } from "./sessions.js"; import { carrySession } from "./session-carry.js"; import { buildEngineMcpRuntime } from "./run-engine-mcp.js"; import { releaseRemoteMcps } from "../mcp/remote-mcp-registry.js"; @@ -361,8 +361,8 @@ export function buildHealedPrompt(ctx, turnText) { // thread's context). Only an EXPLICIT ask for this thread/run — a "claude"/"codex" directive, the // /model wizard's thread scope, or a per-run API engine override — switches it; the caller then // starts the fresh session (and replays the thread transcript into it, like a session heal). The -// usage-limit Codex fallback is separate and untouched: it runs under a suffixed session key and -// never re-stamps the thread. New or unlabeled (pre-v4) sessions have nothing to stick to. +// A successful automatic failover re-stamps the thread with the answering engine, so it too +// continues on that engine. New or unlabeled (pre-v4) sessions have nothing to stick to. export function decideThreadEngine({ requested, sessionEngine, isNew, explicit }) { if (isNew || !sessionEngine || sessionEngine === requested) return { engine: requested, switch: false }; return explicit ? { engine: requested, switch: true } : { engine: sessionEngine, switch: false }; @@ -879,6 +879,17 @@ export async function runMessage({ channelId, authorId, workspaceId = "", text, // tombstone — this run's late unwind (a subprocess that finished at the kill, a resume-heal // re-mint, the Codex-fallback save) is silently dropped instead of resurrecting the session. const sessionGen = sessionGeneration(entry.slug, threadKey); + // Adopt successful fallback sessions written by older gateway versions. Those versions kept + // the answer under a suffixed key and retried the limited engine on the next message. A user's + // explicit thread/runtime choice still wins, including a thread-scoped model selection. + if (!presetSessionId && !engineExplicit) { + const previousEngine = await getSessionEngine(entry.slug, threadKey); + const legacyFallback = fallbackTargets(previousEngine).find(isEngineEnabled); + const pinnedModel = await getThreadModel(entry.slug, threadKey); + if (legacyFallback && (!pinnedModel || !modelBelongsToEngine(pinnedModel, previousEngine))) { + await promoteFallbackSession(entry.slug, threadKey, legacyFallback, "", sessionGen); + } + } if (presetSessionId) { await saveSession(entry.slug, threadKey, presetSessionId, engine, sessionGen, runtimeStamp); sessionId = presetSessionId; @@ -1510,9 +1521,8 @@ export async function runMessage({ channelId, authorId, workspaceId = "", text, // Cross-engine failover for when the engine driving this turn is rate-limited or can't // authenticate: it can't resume the failed engine's session, but its OWN thread is durable — the - // fallback thread id is kept under a suffixed session key, so consecutive fallback turns resume - // the same conversation instead of starting context-less each time (the primary key keeps - // mapping to the original engine's session for when the cooldown ends). Fallback engine + // fallback thread id is initially kept under a suffixed session key while this turn runs; + // after a successful answer it replaces the main thread session. Fallback engine // unavailable → caller handles the thrown error. // Direction-agnostic: `engine` is whatever actually drives this turn, and the target is the first // ENABLED engine in its failover route — an admin who turned a harness off must not have it @@ -1684,7 +1694,7 @@ export async function runMessage({ channelId, authorId, workspaceId = "", text, } } assertCompletedTurn(cx, fallbackEngine, prior || ""); - if (cx.sessionId) await saveSession(entry.slug, fbKey, cx.sessionId, fallbackEngine, sessionGen, runtimeStamp); + if (cx.sessionId) await promoteFallbackSession(entry.slug, threadKey, fallbackEngine, cx.sessionId, sessionGen, runtimeStamp); cx = transientRetryNote(fallbackEngine, cx); // `fellBack`/`fallbackFrom` are the engine-agnostic truth; `fellBackToCodex` is the original // Claude→Codex-only flag, still emitted so existing consumers keep working. @@ -2034,6 +2044,12 @@ export async function runMessage({ channelId, authorId, workspaceId = "", text, if (sessionEngine === "" && !fresh && finalSessionId) { await saveSession(entry.slug, threadKey, finalSessionId, engine, sessionGen, runtimeStamp); } + // A completed turn on the chosen engine supersedes any old fallback-only session. In + // particular, this prevents a legacy fallback from being adopted after a user clears a + // thread-scoped model choice that kept this turn on the original engine. + if (sessionGeneration(entry.slug, threadKey) === sessionGen) { + await discardFallbackSessions(entry.slug, threadKey); + } const finalResult = { ...baseMeta, diff --git a/src/gateway/sessions.js b/src/gateway/sessions.js index d09f067c..3cc3eede 100644 --- a/src/gateway/sessions.js +++ b/src/gateway/sessions.js @@ -50,6 +50,15 @@ function getRow(slug, threadKey) { return getDb().prepare("SELECT session_id, engine, runtime FROM sessions WHERE slug = ? AND thread_key = ?").get(slug, threadKey) || null; } +function deleteFallbackRows(slug, threadKey) { + const prefix = String(threadKey).replace(/[\\%_]/g, (c) => `\\${c}`); + getDb().prepare("DELETE FROM sessions WHERE slug = ? AND thread_key LIKE ? ESCAPE '\\'").run(slug, `${prefix}::%-fallback`); +} + +export async function discardFallbackSessions(slug, threadKey) { + deleteFallbackRows(slug, threadKey); +} + // `engine` records which harness minted `session_id` (a session id is engine-specific — Claude // mints a UUID, Codex mints its own thread_id, and neither can resume the other's). Stamped on // every session write so a later engine switch is detectable (see run.js's harness-switch reset). @@ -118,6 +127,29 @@ export async function saveSession(slug, threadKey, sessionId, engine = "", gener put(slug, threadKey, sessionId, engine, runtime); } +// A successful cross-engine answer becomes the thread's live session. Older gateway versions +// kept that answer only under a fallback key; with no sessionId this also adopts those existing +// rows on the next turn. Keep the clear-generation check and both writes in one synchronous DB +// transaction so /clear cannot leave the main key pointing at a discarded fallback. +export async function promoteFallbackSession(slug, threadKey, engine, sessionId = "", generation = null, runtime = "") { + if (generationStale(slug, threadKey, generation)) return false; + const fallbackKey = `${threadKey}::${engine}-fallback`; + const db = getDb(); + const fallback = getRow(slug, fallbackKey); + const id = sessionId || (fallback?.engine === engine ? fallback.session_id : ""); + if (!id) return false; + db.exec("BEGIN IMMEDIATE"); + try { + put(slug, threadKey, id, engine, sessionId ? runtime : fallback.runtime); + deleteFallbackRows(slug, threadKey); + db.exec("COMMIT"); + return true; + } catch (error) { + db.exec("ROLLBACK"); + throw error; + } +} + export async function getSessionMap(slug) { const rows = getDb().prepare("SELECT thread_key, session_id FROM sessions WHERE slug = ?").all(slug); const map = {}; @@ -158,10 +190,7 @@ export async function clearSession(slug, threadKey) { .run(slug, threadKey); // LIKE with an explicit ESCAPE so a thread key containing % or _ can't widen the delete into // other threads' rows. The trailing "-fallback" keeps it to fallback keys only. - const prefix = String(threadKey).replace(/[\\%_]/g, (c) => `\\${c}`); - getDb() - .prepare("DELETE FROM sessions WHERE slug = ? AND thread_key LIKE ? ESCAPE '\\'") - .run(slug, `${prefix}::%-fallback`); + deleteFallbackRows(slug, threadKey); return info.changes > 0; } diff --git a/test/claude-fallback-e2e.test.js b/test/claude-fallback-e2e.test.js index 3ecbbcb8..901456d7 100644 --- a/test/claude-fallback-e2e.test.js +++ b/test/claude-fallback-e2e.test.js @@ -16,6 +16,7 @@ const { setUser, upsertChannelEntry, saveChannelMeta } = await import("../src/co const { saveSettings } = await import("../src/config/settings.js"); const { runMessage, resetEngineCooldowns } = await import("../src/gateway/run.js"); const { setThreadModel } = await import("../src/gateway/thread-engine.js"); +const { getSession, getSessionEngine, saveSession } = await import("../src/gateway/sessions.js"); const claudeDM = async (id, name) => { const entry = await upsertChannelEntry(id, { name, type: "im", isDM: true }); @@ -55,6 +56,54 @@ test("a replay-safe thrown Claude usage-limit failure transparently falls back t assert.equal(result.fellBackToCodex, true); assert.match(result.content, /usage limit.*using Codex/i); assert.match(result.content, /Codex stub reply/); + assert.equal(await getSessionEngine(entry.slug, "1900.050"), "codex"); + resetEngineCooldowns(); + const next = await runMessage({ + channelId: "D_SAFE_LIMIT", authorId: "U_SAFE_LIMIT", text: "continue", + threadKey: "1900.050", origin: "slack_foreground", preferCold: true, + }); + assert.equal(next.engine, "codex", "the successful fallback remains the thread's engine after cooldown"); + assert.match(next.content, /resume=yes/); + assert.doesNotMatch(next.content, /using Codex/i, "a normal continuation is not another fallback"); +}); + +test("a successful fallback written by an older gateway is adopted before retrying Claude", async () => { + resetEngineCooldowns(); + saveSettings({ engine: "claude", codexFallback: true, composioMode: "personal" }); + await setUser("U_LEGACY_FALLBACK", { name: "Legacy Fallback", approved: true, isAdmin: false }); + const entry = await claudeDM("D_LEGACY_FALLBACK", "legacy-fallback"); + await saveSession(entry.slug, "1900.055", "old-claude-session", "claude"); + await saveSession(entry.slug, "1900.055::codex-fallback", "codex-stub-thread", "codex"); + const result = await runMessage({ + channelId: "D_LEGACY_FALLBACK", authorId: "U_LEGACY_FALLBACK", text: "continue", + threadKey: "1900.055", origin: "slack_foreground", preferCold: true, + }); + assert.equal(result.engine, "codex"); + assert.match(result.content, /resume=yes/); + assert.equal(await getSessionEngine(entry.slug, "1900.055"), "codex"); + assert.equal(await getSession(entry.slug, "1900.055::codex-fallback"), null); +}); + +test("a manual model choice keeps the original engine and retires an old fallback row", async () => { + resetEngineCooldowns(); + saveSettings({ engine: "claude", codexFallback: true, composioMode: "personal" }); + await setUser("U_LEGACY_PIN", { name: "Legacy Pin", approved: true, isAdmin: false }); + const entry = await claudeDM("D_LEGACY_PIN", "legacy-pin"); + await saveSession(entry.slug, "1900.056", "stub-session", "claude"); + await saveSession(entry.slug, "1900.056::codex-fallback", "codex-stub-thread", "codex"); + await setThreadModel(entry.slug, "1900.056", "opus"); + const result = await runMessage({ + channelId: "D_LEGACY_PIN", authorId: "U_LEGACY_PIN", text: "continue", + threadKey: "1900.056", origin: "slack_foreground", preferCold: true, + }); + assert.equal(result.engine, "claude"); + assert.equal(await getSession(entry.slug, "1900.056::codex-fallback"), null); + await setThreadModel(entry.slug, "1900.056", ""); + const next = await runMessage({ + channelId: "D_LEGACY_PIN", authorId: "U_LEGACY_PIN", text: "continue again", + threadKey: "1900.056", origin: "slack_foreground", preferCold: true, + }); + assert.equal(next.engine, "claude"); }); test("a Codex cross-engine fallback also replaces a rejected channel model with its gateway default", async () => { diff --git a/test/codex-failover-e2e.test.js b/test/codex-failover-e2e.test.js index 47a6cc15..62b6f842 100644 --- a/test/codex-failover-e2e.test.js +++ b/test/codex-failover-e2e.test.js @@ -22,6 +22,7 @@ process.env.SESSION_KEEPALIVE = "0"; const { setUser, upsertChannelEntry, saveChannelMeta } = await import("../src/config/store.js"); const { saveSettings } = await import("../src/config/settings.js"); const { runMessage, resetEngineCooldowns } = await import("../src/gateway/run.js"); +const { getSessionEngine } = await import("../src/gateway/sessions.js"); const { runCodex } = await import("../src/engines/codex.js"); const { setThreadEngine, setThreadModel } = await import("../src/gateway/thread-engine.js"); @@ -38,7 +39,7 @@ test("a Codex usage limit reported as a JSON error event falls back to Claude", resetEngineCooldowns(); saveSettings({ engine: "codex", engineFallback: true, engineEnabled: { claude: true, codex: true }, composioMode: "personal" }); await setUser("U_CODEX_LIMIT", { name: "Codex Limit", approved: true, isAdmin: false }); - await codexChannel("D_CODEX_LIMIT", "codex-limit"); + const entry = await codexChannel("D_CODEX_LIMIT", "codex-limit"); const result = await runMessage({ channelId: "D_CODEX_LIMIT", @@ -56,6 +57,14 @@ test("a Codex usage limit reported as a JSON error event falls back to Claude", assert.equal(result.fellBackToCodex, false, "the legacy flag names the TARGET, and Claude is not Codex"); assert.match(result.content, /Codex hit its usage limit.*using Claude/i); assert.match(result.content, /Stub engine reply/); + assert.equal(await getSessionEngine(entry.slug, "1901.010"), "claude"); + resetEngineCooldowns(); + const next = await runMessage({ + channelId: "D_CODEX_LIMIT", authorId: "U_CODEX_LIMIT", text: "continue", + threadKey: "1901.010", origin: "slack_foreground", preferCold: true, + }); + assert.equal(next.engine, "claude"); + assert.match(next.content, /resume=yes/); }); test("the same limit printed on stderr with a nonzero exit also falls back", async () => { diff --git a/test/runtime-identity-preamble.test.js b/test/runtime-identity-preamble.test.js index 393820aa..ed322b11 100644 --- a/test/runtime-identity-preamble.test.js +++ b/test/runtime-identity-preamble.test.js @@ -220,9 +220,9 @@ test("cross-engine fallback and its model retry use their own engine, model, eff const followupStart = backend.calls.spawn.length; await runMessage({ ...context, text: "CODEX_STUB_REJECT_MODEL" }); const followup = attempts(followupStart); - assert.equal(followup.length, 2, "the cooldown resumes the existing fallback conversation"); + assert.equal(followup.length, 2, "the promoted Codex session resumes after its model retry"); for (const call of followup) assert.ok(call.args.some((arg) => String(arg).includes("Clean mode for this attempt: **enabled**"))); - assert.deepEqual(metadata(followup[1]), { engine: "codex", configured_model: "gpt-5.6-sol", configured_effort: null, session: "resumed" }); + assert.deepEqual(metadata(followup[1]), { engine: "codex", configured_model: "gpt-5.6-sol", configured_effort: "high", session: "resumed" }); }); test("a lost resumed session gets fresh metadata when the gateway heals it", async (t) => { From 4bd1c4b697e5770bc452a29c9938ee591e0eacbf Mon Sep 17 00:00:00 2001 From: Tiberiu Socaci Date: Tue, 29 Sep 2026 15:58:26 +0300 Subject: [PATCH 06/16] feat: allow channel-specific Codex authentication Signed-off-by: Tiberiu Socaci --- CHANGELOG.md | 3 ++ FEATURES.md | 12 +++++ TEST-PLAN.md | 24 ++++++++++ docs/ENGINE-CAPABILITIES.md | 2 +- docs/OPERATIONS.md | 46 +++++++++++++------ docs/SSH-ACCESS.md | 5 +- public/app.js | 3 ++ public/index.html | 9 ++++ src/config/channel-audit.js | 1 + src/config/store.js | 1 + src/engines/codex.js | 20 ++++++-- src/gateway/channel-codex-auth.js | 21 +++++++++ src/gateway/codex-token-relay.js | 23 ++++++++-- src/gateway/egress/catalog-rules.js | 4 +- src/gateway/egress/grants.js | 42 ++++++++++++----- .../references/credentials-connections.md | 17 +++---- src/runtimes/container/credentials.js | 42 ++++++++++------- src/runtimes/container/lifecycle.js | 2 +- src/web/routes/channels.js | 4 ++ test/codex-token-relay.test.js | 42 ++++++++++++++++- test/container-credentials.test.js | 35 ++++++++++++-- 21 files changed, 292 insertions(+), 66 deletions(-) create mode 100644 src/gateway/channel-codex-auth.js diff --git a/CHANGELOG.md b/CHANGELOG.md index 6a1d03b7..bc8fca68 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -18,6 +18,9 @@ product overview. ## Unreleased +- Channels can select a separate host-side Codex login for ChatGPT subscription access or an + OpenAI API key. Proxy-mode containers receive only channel-bound credential placeholders. + - **Automatic engine switches now stay with the conversation.** When Claude reaches its limit and Codex answers, the next message resumes that Codex session instead of retrying Claude. The same holds when Codex switches to Claude. Existing threads with a successful fallback session adopt diff --git a/FEATURES.md b/FEATURES.md index 9f2a9d23..0fc26389 100644 --- a/FEATURES.md +++ b/FEATURES.md @@ -1,5 +1,17 @@ # ChannelGate — Features +## Channel-specific Codex authentication + +- Admins can select the gateway Codex login or a dedicated host-side `CODEX_HOME` for a channel + from Runtime → Codex authentication. The dedicated directory is keyed to the conversation ID + and does not fall back to another account when no login exists. +- The operator can sign in to that directory with ChatGPT subscription access or an OpenAI + Platform API key. In proxy-mode containers, both methods use a channel-bound placeholder: + ChatGPT access tokens are relayed as JWT-shaped values, while API keys are relayed only to + `api.openai.com`. The host `auth.json` and any refresh token stay outside the container. +- The selected login also applies to an admin's direct-host Codex turn. The authentication + source is recorded in channel policy audits; no credential value is included. + ## Optional isolated VPN database service - The host operator can provision a per-channel OpenVPN 3 Linux service and an unprivileged MySQL diff --git a/TEST-PLAN.md b/TEST-PLAN.md index a3af8f81..54e6bd37 100644 --- a/TEST-PLAN.md +++ b/TEST-PLAN.md @@ -1,5 +1,29 @@ # ChannelGate — Test Plan +## Channel-specific Codex authentication (2026-09-29) + +- [x] Automated: `node --test test/codex-token-relay.test.js + test/container-credentials.test.js test/engine-runtime-isolated.test.js + test/codex-args.test.js test/channel-env.test.js`. A channel-selected login has no gateway + fallback. A synthetic subscription cache yields an access-only channel placeholder and leaves + the refresh token on the host. A synthetic API-key cache yields a separate placeholder swapped + only in the Authorization header at `api.openai.com`; the real key never enters container + `auth.json`. Proxy mode never mounts either host login file. +- [x] CLI shape check: Codex CLI 0.156.1 `login status` accepts a temporary, synthetic API-key + `auth.json` with `auth_mode: "apikey"`; no real credential or provider request was used. +- [ ] Live Codex acceptance: in a disposable channel, select **This channel's login**, sign in + as a different permitted ChatGPT account under the displayed host `CODEX_HOME`, and ask for a + harmless answer. Pass: Codex answers, the host file keeps the refresh token, the container file + has only the channel's placeholder and an empty refresh token, and the egress audit records a + relay on the Codex hosts. Then remove the channel file while the gateway login remains valid: + the channel must fail authentication rather than use the gateway account. Restore afterward. +- [ ] Live Codex API-key acceptance: use a disposable OpenAI project key in a second channel's + displayed host `CODEX_HOME`, send the same harmless prompt, and require a response and an + `api.openai.com` relay audit entry. The container file contains only the placeholder; a copied + placeholder sent to `chatgpt.com` is not swapped. Revoke the test key afterward. +- [ ] Live Claude isolation: with either Codex channel choice selected, run a Claude turn in the + same disposable channel and require the existing Claude login and normal reply. + ## Live-case definitions corrected for the container-secrets contract (2026-09-27 QA campaign) The 2026-09-27 live campaign failed or blocked these registry cases only because their written diff --git a/docs/ENGINE-CAPABILITIES.md b/docs/ENGINE-CAPABILITIES.md index 73444143..efaee2d4 100644 --- a/docs/ENGINE-CAPABILITIES.md +++ b/docs/ENGINE-CAPABILITIES.md @@ -15,7 +15,7 @@ the Admin API/UI consume that registry. | Skills | Organization/channel repository skills plus per-author grants | Native organization/channel repository skills plus a per-run personal skill catalog; personal delivery does not register slash commands | Same as Claude (`CLAUDE.md`, `.claude/skills`, plugin dirs) | | Usage/cost | provider-reported cost | token usage with configured rate estimate | token usage only — the CLI's Anthropic-priced figure is dropped and no rate is inferred | | Health | adapter-owned `--version` boot probe | adapter-owned `--version` boot probe | adapter-owned `--version` boot probe plus "is a QwenCloud key configured" | -| Container login | a relay of the host's Claude access token in `CLAUDE_CODE_OAUTH_TOKEN` — behind the egress proxy the channel's `sk-ant-oat01-cgph_r…` placeholder, swapped on `api.anthropic.com`; refreshed by a cheap host turn (`src/gateway/claude-token-relay.js`); never a file | behind the egress proxy an ACCESS-ONLY `auth.json` in the channel HOME, written before each run: a JWT-shaped `cgph_r…` placeholder swapped whole on `api.openai.com`, `chatgpt.com`, `auth.openai.com`, empty refresh token; refreshed by a cheap ephemeral host turn (`src/gateway/codex-token-relay.js`). Legacy bridge mode or an API-key login: the shared read-write file mount | the gateway's provider key in `ANTHROPIC_AUTH_TOKEN`, raw (not relayed) | +| Container login | a relay of the host's Claude access token in `CLAUDE_CODE_OAUTH_TOKEN` — behind the egress proxy the channel's `sk-ant-oat01-cgph_r…` placeholder, swapped on `api.anthropic.com`; refreshed by a cheap host turn (`src/gateway/claude-token-relay.js`); never a file | behind the egress proxy an ACCESS-ONLY `auth.json` in the channel HOME, written before each run: a JWT-shaped placeholder for a ChatGPT login or an API-key placeholder restricted to `api.openai.com`; no refresh token or real API key in the container. The host login can be gateway-wide or channel-specific. Legacy bridge mode keeps a read-write file mount | the gateway's provider key in `ANTHROPIC_AUTH_TOKEN`, raw (not relayed) | The Qwen column covers every harness generated from the Anthropic-compatible **provider table** in `src/engines/qwen.js` — today `qwen` (QwenCloud) and `qwen-eu` (Alibaba Cloud Model Studio, EU diff --git a/docs/OPERATIONS.md b/docs/OPERATIONS.md index 219f16ce..644ccaa5 100644 --- a/docs/OPERATIONS.md +++ b/docs/OPERATIONS.md @@ -372,8 +372,8 @@ in on the host is all a channel needs. Nothing is copied or mounted. Optionally `claude setup-token` on the gateway host and paste the value into *Claude token for container runs*: that token is then used instead and never needs refreshing. Codex is RELAYED the same way behind the egress proxy (see **Codex login relay** below): keep the host signed in with `codex login`; -nothing is mounted. Only the legacy open-network mode and an API-key Codex login still bind-mount -the gateway's real `auth.json` read-write into every container (Codex rewrites it in place, so a +nothing is mounted. Only the legacy open-network mode still bind-mounts +the selected `auth.json` read-write into a container (Codex rewrites it in place, so a copy would fork the refresh chain). Codex *sessions* and history are per channel either way. **Network (the egress proxy).** Every channel container runs with `--network none`: its only @@ -650,7 +650,7 @@ gone — a finished test run, not a second live gateway sharing this account. A is never removed by it. The space estimate counts image layers once each. To make it routine, schedule the report and read it; schedule `--apply` only if you have decided to. -Codex uses the same shared login file already mounted for chat turns. Claude's rotating credential +Codex uses the channel's selected login, relayed in proxy mode. Claude's rotating credential file is still never copied or mounted: the helper refreshes the gateway's normal subscription access-token relay every 20 minutes and exposes only that access token to interactive `claude` commands. Closing VS Code removes the live token and releases the lease; an interrupted helper is @@ -696,9 +696,9 @@ per-channel or gateway-wide "back to the host" switch: the container is the only - **Per-user Codex skill grants are not delivered in containers.** The per-run Codex skill overlay was built for a synthetic host HOME that a container does not have; a Codex run gets the channel's skills through the mounted workdir, but not that overlay. -- **Codex sessions are per channel; the sign-in is shared only outside the proxy.** Under the - legacy open-network mode (or with an API-key `auth.json`) every container mounts the same - `auth.json` the gateway uses. A `codex login` on the host that *replaces* the file leaves a +- **Codex sessions are per channel; the sign-in is mounted only outside the proxy.** Under the + legacy open-network mode each container mounts its selected host `auth.json`. A `codex login` + on the host that *replaces* the file leaves a running container holding the old inode — `/status` and `/api/health` report the drift; restart the channel's container (or let the reaper stop it) to pick the new one up. Behind the proxy the login is relayed and a new `codex login` takes effect on the next turn. @@ -709,10 +709,8 @@ per-channel or gateway-wide "back to the host" switch: the container is the only dies within hours). A daemon that authenticates Claude with its own `ANTHROPIC_API_KEY` still passes that key through raw — the proxy does not rewrite it. - **Engine keys that still reach a proxy-mode container raw.** A Qwen harness's provider key - (`ANTHROPIC_AUTH_TOKEN` pointed at the provider) and a daemon `CODEX_API_KEY`/`OPENAI_API_KEY` - handed to Codex are real values in the container environment; Codex's own sign-in is the shared - `auth.json` mount (the engine-login broker is a later phase). They are reachable only on their - engine endpoints through the proxy, but a process in the container can read them. + (`ANTHROPIC_AUTH_TOKEN` pointed at the provider) still reaches that engine as a real value. + Codex logins stored in `auth.json`, including API-key logins, are relayed behind the proxy. - **A self-hosted Qwen endpoint on a private address is refused in proxy mode.** The proxy never connects to loopback, private, link-local or CGNAT addresses, and a configured Qwen base URL is no exception: point the harness at a public endpoint, or run that channel under the legacy bridge @@ -727,9 +725,8 @@ per-channel or gateway-wide "back to the host" switch: the container is the only the names of any unprotected (`egressUnprotected`) or withheld (`egressWithheld`) secrets. `networkEnforcedFor(target)` in `src/engines/network-policy.js` is the one question every surface asks. -- **Codex's sign-in is a placeholder in proxy mode.** See **Codex login relay** below. What remains - raw: the legacy bridge mode's shared file, an API-key `auth.json`, and a daemon - `OPENAI_API_KEY`/`CODEX_API_KEY`, which reaches the container env as `CODEX_API_KEY` unchanged. +- **Codex's sign-in is a placeholder in proxy mode.** See **Codex login relay** below. The legacy + bridge mode still mounts the selected `auth.json` file because it has no proxy swap. - **SSH and VS Code sessions** run in the same `--network none` container with the proxy env and hold placeholders like a turn (container-secrets P3, `docs/SSH-ACCESS.md`); SSH `-L` forwards to external hosts do not work without the network. @@ -755,6 +752,29 @@ turn after upgrading recreates each channel's container once (the mount is gone fingerprint), and `cg-init` deletes an old copied Codex login carrying a refresh token from the volume. The legacy open-network mode keeps the shared read-write mount (and says so in `/status`). +**A separate Codex login for one channel.** In the admin channel editor, Runtime → **Codex +authentication** → **This channel's login**, then Save. The editor shows that channel's Codex +home directory on the **gateway host**. Create it under the gateway service account and sign in +there. For ChatGPT subscription access, use `CODEX_HOME= codex login +--device-auth` and complete the browser code flow. For an OpenAI Platform API key, use +`CODEX_HOME= codex login --with-api-key`, supplying the key on standard +input as prompted by the CLI. Never put the key on the command line, in Slack, or in the channel +folder. `CODEX_HOME= codex login status` checks the selected method. OpenAI +bills API-key runs separately from a ChatGPT subscription. The channel login does not fall back +to the gateway login if it is missing or expired. + +For a laptop login, OpenAI documents copying `~/.codex/auth.json` to a headless host as a +fallback when device login is unavailable. Use an encrypted SSH transfer into the **displayed +host directory** with restrictive permissions and run Codex as the gateway service account so +the one host-side file can refresh. The file contains full credentials, including a refresh +token: do not upload it through Slack or the admin UI. Channel containers receive only a +channel-bound placeholder. A ChatGPT login is relayed as the JWT-shaped access token described +above; an API-key login is relayed as a separate placeholder swapped only in the Authorization +header at `api.openai.com`. The channel setting affects new Codex turns; it does not change +Claude authentication. An Admin channel with the optional full operator-home mount can read +host files under that mount, including the gateway's runtime root; use ordinary project channels +for credential isolation. + **Remote MCP relay (container-secrets P1).** Composio (`composio-user`, `composio-agent` in token mode), the MakeItFuture toolbox and the Make toolbox never reach a container with their token. The engine's MCP entry is the image's socket bridge naming the server (`CG_MCP_SERVICE=remote-mcp`); diff --git a/docs/SSH-ACCESS.md b/docs/SSH-ACCESS.md index f173fda7..b2af4a26 100644 --- a/docs/SSH-ACCESS.md +++ b/docs/SSH-ACCESS.md @@ -102,7 +102,8 @@ machine settings — merge-only, and never over a value you set yourself. Inside, you are user `agent` in the channel's work folder (an interactive login starts there; image spec 1.5.1), with the same environment an engine turn gets and the channel's persistent `/home/agent` (installed tools, `gh`/`vercel`/`supabase` logins, Claude and Codex history). -Codex uses the shared sign-in mount, and `codex` in a session gets what a chat turn's Codex gets: +Codex uses the channel's selected login (a proxy placeholder in proxy mode, a host file mount in +legacy bridge mode), and `codex` in a session gets what a chat turn's Codex gets: the gateway tools, your own Composio accounts as `composio-user`, the channel's as `composio-agent`, the channel's selected MCP servers and your channel secrets. Started from your home, `/` or a parent of the channel folder it moves into the channel folder so its `AGENTS.md` @@ -146,7 +147,7 @@ session gets: expiry and rate tier beside it are the real login's — they are facts, not secrets), so the file is useless outside the container. `claude -r ` resumes a thread's own session. -Codex over SSH is unchanged (its login is the shared sign-in mount). "show SSH access" in the +Codex over SSH uses the same selected login as a channel turn. "show SSH access" in the channel prints what a session gets. If a part could not be prepared — no Claude login on the host, a channel whose Composio session is unavailable — the attach still succeeds and the daemon log names the part. diff --git a/public/app.js b/public/app.js index 21c97c54..5d70e59f 100644 --- a/public/app.js +++ b/public/app.js @@ -2060,6 +2060,8 @@ function renderChannelDetail(ch) { } const engineSelect = card.querySelector(".ch-engine"); engineSelect.value = meta.engine || ""; + card.querySelector(".ch-codex-auth-source").value = meta.codexAuthSource || "gateway"; + card.querySelector(".ch-codex-auth-home").textContent = meta.codexAuthHome || "Save this channel first"; renderMcpBoxForEngine(mcpsBox, engineSelect.value, mcpsCount); syncModelOptions({ engineSelect, @@ -2191,6 +2193,7 @@ function renderChannelDetail(ch) { memory: card.querySelector(".ch-memory").checked, noDefaultTokens: card.querySelector(".ch-nodefaulttokens").checked, engine: engineSelect.value, + codexAuthSource: card.querySelector(".ch-codex-auth-source").value, workDir: card.querySelector(".ch-workdir").value, syncDriveFolder: card.querySelector(".ch-syncdrive").value, model: card.querySelector(".ch-model").value, diff --git a/public/index.html b/public/index.html index 6bbdb7dc..6640d8f4 100644 --- a/public/index.html +++ b/public/index.html @@ -1273,6 +1273,15 @@

Environment variables

+ +
Channel Codex home on gateway host
Channel Codex home on gateway host
+