Skip to content

feat: GET /v1/usage answers for the key or its owner; API key limits in the console - #93

Merged
fylorn merged 4 commits into
devfrom
feat/key-usage-endpoint
Oct 10, 2026
Merged

fylorn merged 4 commits into
devfrom
feat/key-usage-endpoint

Conversation

@fylorn

@fylorn fylorn commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Two things, both decided for this release:

  1. GET /v1/usage on the gateway port tells a client holding a gateway API key what room it has left, about one subject, named in a top-level scope.
  2. The console edits an API key's own limits. The key's edit dialog gets a Limits tab for the rate limits and budgets the backend already stores on a key lineage and enforces.

GET /v1/usage: the rule

  • "scope": "key" when the key has any limit of its own on the AI gateway (a rate-limit rule or a budget on its lineage). limits lists only the key's limits, each with used counted for the key; usage is the key's own (across its rotations); cost_usd_month is the key's.
  • "scope": "user" when the key has none. limits lists the owner's effective limits (roles, including team-granted roles, merged most-restrictive, then the user's overrides), each with used counted for everything the owner does; usage is the owner's total over all of their keys; cost_usd_month is summed over all of their keys.
  • Neither the key nor the owner has a limit: "scope": "user", "limits": [], the owner's totals.

scope is always present and is exactly "key" or "user"; each limit entry keeps its own scope too (always equal to the top-level one). A key whose only limit is on the MCP gateway answers for its owner.

{
  "scope": "key",
  "usage": {"requests_today": 12, "tokens_today": 48210,
            "requests_month": 340, "tokens_month": 1290455,
            "cost_usd_month": 3.82},
  "limits": [
    {"scope": "key", "kind": "tokens", "window": "daily", "window_secs": null,
     "limit": 100000, "used": 48210, "resets_at": "2026-10-12T00:00:00Z"},
    {"scope": "key", "kind": "requests", "window": "1m", "window_secs": 60,
     "limit": 60, "used": 4, "resets_at": null}
  ],
  "expires_at": "2026-12-31T00:00:00Z"
}

A holder of a key without limits of its own sees the account's totals: the owner's requests, weighted tokens and cost today/this month across every key, and the owner's limits with what the whole account has used of them. A key handed to someone else should carry limits of its own if the owner's numbers are not theirs to see.

  • limits is built with limits_for_ai_gateway (what the pre-flight stages enforce) and then filtered to the chosen subject. used is read from the same Redis counters the same way (sliding::read_count, budget::read_spend), without charging them. Ordered by the share left, least first; ties put shorter sliding windows first, then calendar periods, then requests before tokens. MCP-surface limits are not listed.
  • usage comes from per-key and per-user day/month counters (limits::usage, renamed from the earlier key_usage). The user-subject counters that exist are limit counters: kept only for configured limits and only since they were configured, so they cannot give a day or month total for a user, and new per-user counters sit next to the per-key ones. Every counted request writes four hashes under the owner's hash tag in one script:
    usage:{user:<owner>}:api_key_lineage:<lineage>:daily:<date> / :monthly:<month> and usage:{user:<owner>}:user:<owner>:daily:<date> / :monthly:<month>, TTL 2× the period. They count at the same points as before: a request once it passes the rate limits, its weighted tokens in post_flight_account. Answers from the response cache now count their tokens (the stored prompt + completion totals, weighted for the model as a call's are), so usage agrees with limits and budgets once those count cache hits too (see "Merging with fix/limit-enforcement"). Writes fail open with gateway_usage_count_fail_open_total.
  • cost_usd_month: sum(cost_usd) from gateway_logs since the 1st (UTC), by api_key_lineage_id for a key and by user_id for a user, as the cost reports sum it; null without ClickHouse or on a failed read.
  • expires_at, errors (same 401/403 as a model request, 503 when the counters can't be read) and the read-only middleware (no last_used_at update, no early-cancel request row, nothing charged) are unchanged from the first commit.

Key limits in the console

  • Edit API key → Limits tab: the key's rules (requests or weighted tokens over 1m, 5m, 1h, 5h, 1d, 1w) and budgets (daily, weekly, monthly weighted tokens), each with a live usage meter, through the existing endpoints: GET /api/admin/limits/api_key/{id}/rules|budgets|usage, POST …/rules|budgets, DELETE …/rules/{id}|budgets/{id}. No new endpoint, no backend permission change.
  • Reused, not copied: the usage meter, expiry cell and rule label are exported from the user limits tab, and its override drawer becomes LimitDrawer, which takes the subject (user or api_key). For a key it defaults to no expiry and, in an edit, prefills the row's expiry and reason.
  • Edits keep the slot: in an edit the type, gateway, metric, window and period selects are locked, for users and keys alike. Rows are upserted by that slot, so changing one used to create a second row rather than edit the first. This also changes the user page's "Edit" on an override.
  • Permissions as the endpoints decide: the tab shows to holders of rate_limits:read, and disappears if the endpoints answer 403 (e.g. a team-scoped grant, which never covers the caller's own keys). Without rate_limits:write it is read-only (no add, edit or delete). Everyone else sees the dialog exactly as before.
  • i18n: en + zh (pnpm check:i18n parity holds). The Configuration Guide line for /v1/usage now states the rule.

Security-relevant

  • Authentication is the existing API-key middleware body, unchanged by this round. The response carries no ids, names or emails.
  • What a key holder learns is now explicit and depends on the key: with limits of its own, only the key's limits and the key's own usage and cost; without, the owner's limits and the owner's account-wide usage and cost across all of their keys. The owner's per-key breakdown is not exposed.
  • The console tab adds no endpoint and widens no permission: every call goes through the existing limits handlers (rate_limits:read|write + assert_scope_for_subject).

Tests

  • Unit (handlers::key_usage): neither has limits → user, []; key without limits (an MCP-only key rule included) → owner's limits only; key with limits while the owner has some → key's limits only; a key budget alone makes the answer the key's; ordering and window labels. limits::usage: the four hashes share one cluster slot; day/month turnover.
  • Integration (limits.rs):
    • key with limits → scope: key, only the key's limits and usage, although the owner has limits (one with less left) and another key of the owner is used;
    • keys without limits → scope: user, the owner's role / team-role / override limits with used over both keys, and usage the owner's totals including the other key's requests; both keys get the same answer;
    • neither → scope: user, limits: [], totals over two keys (via x-api-key);
    • an answer from the cache counts its request and tokens, on the key (scope: key) and on the owner (scope: user from another key);
    • 401/403 identical to a model request; calling it moves no counter and no last_used_at.
  • Integration (analytics_clickhouse.rs): cost_usd_month is the key's for a key with limits (2 calls) and the owner's for a key without (3 calls).
  • Web: key-limits-tab.test.tsx (lists rules and budgets with usage; empty state; read-only without write; add posts a permanent rule to api_key/{id}; edit keeps the slot, expiry and reason; delete a budget after confirmation) and api-keys.test.tsx (the tab in the edit dialog; no tab without rate_limits:read; no tab on 403).
cargo fmt --all -- --check
cargo clippy --workspace --all-targets -- -D warnings
cargo nextest run --workspace --lib --bins                                        # 789 passed
cargo nextest run -p think-watch-test-support --run-ignored only --profile ci     # 426 passed (dedicated PG/Redis/CH containers)
cd web && pnpm check:i18n && pnpm lint && pnpm test && pnpm build                 # 152 tests passed

Not covered by a test: the 503 path (the harness has no Redis fault injection). The new tab has not been checked by eye in a browser.

Merging with fix/limit-enforcement

That branch makes cache hits call post_flight_account, which (here) records the usage tokens. Both branches touch the cache-hit block in proxy/generate.rs and the pre-flight in proxy/pipeline.rs, so whichever lands second gets conflicts there. Resolve by:

  • dropping the count_cache_hit_usage call in the cache-hit block (and the now-unused function in accounting.rs): post_flight_account counts those tokens, and keeping both would count them twice;
  • keeping the usage::record_request call right after check_limits, wherever that stage ends up in the new order (budget → access → limits).

New messages

None (no core message codes involved).

🤖 Generated with Claude Code

fylorn and others added 3 commits October 10, 2026 22:18
…ding it

A client holding a gateway API key had no way to ask how much room the
key has left before a request is refused. GET /v1/usage answers that on
the gateway port, authenticated exactly as a model request is.

`limits` is built with `limits_for_ai_gateway`, the function the
pre-flight stages use, so it lists the same rules and caps a request
is held to: the key's own (scope "key", counted on the lineage) and
the owner's effective limits (scope "user": roles, including
team-granted roles, merged most-restrictive, then overrides; counted
on the owner's counters, shared by all of the owner's keys). `used` is
read from the same Redis counters the same way, without charging
them, and the list is ordered by the share left, least first. A Redis
read error answers 503 rather than zeros.

`usage` needs a count for keys that have no limits, which no existing
counter provides (limit counters exist only for configured limits; the
request log is optional and holds raw tokens). The gateway now counts
each key lineage's requests and weighted tokens per UTC day and month
in Redis, at the points the limits engine counts them: a request once
it passes the rate limits, its tokens in post-flight accounting. The
writes fail open. The month's cost comes from gateway_logs, as the cost
reports read it, and is null without ClickHouse.

The endpoint sits behind a read-only variant of the API-key
middleware: same lookup, checks and refusals, but it neither updates
last_used_at (a poller must not keep an idle key past its inactivity
timeout) nor arms the early-cancel request row.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The answer used to mix the key's own limits with its owner's, next to a
usage count for the key alone, so a client could not tell whose room the
numbers described. It now picks one subject and says which in a
top-level `scope`:

- "key" when the key has limits of its own on the AI gateway: `limits`
  lists only those, counted for the key, and `usage` is the key's own.
- "user" when it has none: `limits` lists the owner's effective limits
  (roles, including team-granted ones, merged most-restrictive, then
  overrides), counted for everything the owner does, and `usage` is the
  owner's total over all of their keys, `cost_usd_month` included. With
  no limits on the owner either, `limits` is empty.

The limits still come from `limits_for_ai_gateway`, filtered to the
chosen subject, and `used` is read as before. A key whose only limit is
on the MCP gateway answers for its owner.

The owner's totals need a counter of their own: the user-subject
counters that exist are limit counters, kept only for configured limits
and only since they were configured, so they cannot give a day or month
total. The per-key counters become per-key and per-user ones in one
module (`limits::usage`, was `key_usage`): every counted request and its
tokens land on the key's and the owner's day and month hashes, all four
under the owner's `{user:<id>}` hash tag, one script per count, at the
same two points as before.

An answer from the response cache now counts its tokens on these
counters too, weighted as a call's are, so the usage agrees with the
limits once those count cache hits as well.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The backend has long held rate-limit rules and budgets on a key lineage
and enforced them, but the console could only set limits on users and
roles. The key's edit dialog now has a Limits tab: the key's rules
(requests or weighted tokens over 1m … 1w) and budgets (daily, weekly,
monthly weighted tokens), each with what it has used, read through the
existing `/api/admin/limits/api_key/{id}/rules|budgets|usage` endpoints
and added, changed and removed through the existing POST and DELETE
ones.

The tab reuses the user limits tab's parts rather than copying them:
the usage meter, expiry cell and rule label are exported from it, and
its override drawer becomes `LimitDrawer`, taking the subject to write
to. For a key the drawer defaults to no expiry and, when editing,
prefills the row's expiry and reason. In an edit the slot fields (type,
gateway, metric, window, period) are locked for both users and keys:
the row is stored by that slot, so changing one created a second row
instead of editing the first.

Permissions are the endpoints' own: the tab shows to holders of
`rate_limits:read` and disappears when the endpoints answer 403 (a
team-scoped grant never covers the caller's own keys); without
`rate_limits:write` it is read-only. Everyone else sees the dialog as
before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@fylorn fylorn changed the title feat(gateway): GET /v1/usage reports a key's usage and the limits binding it feat: GET /v1/usage answers for the key or its owner; API key limits in the console Oct 10, 2026
Resolves the overlap with #94: cached answers are accounted once, through
post_flight_account, which also records the key and user usage counters
(the separate count_cache_hit_usage is gone); usage::record_request runs
after check_limits in the new budget, access, limits order; current_spend
propagates a counter read error. Removes the gateway RateLimiter that was
built into GatewayState and never called.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@fylorn
fylorn merged commit 31cbfc8 into dev Oct 10, 2026
6 checks passed
@fylorn
fylorn deleted the feat/key-usage-endpoint branch October 10, 2026 15:39
@fylorn fylorn mentioned this pull request Oct 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant