Repository navigation
feat: GET /v1/usage answers for the key or its owner; API key limits in the console - #93
Merged
Merged
Conversation
…ding it A client holding a gateway API key had no way to ask how much room the key has left before a request is refused. GET /v1/usage answers that on the gateway port, authenticated exactly as a model request is. `limits` is built with `limits_for_ai_gateway`, the function the pre-flight stages use, so it lists the same rules and caps a request is held to: the key's own (scope "key", counted on the lineage) and the owner's effective limits (scope "user": roles, including team-granted roles, merged most-restrictive, then overrides; counted on the owner's counters, shared by all of the owner's keys). `used` is read from the same Redis counters the same way, without charging them, and the list is ordered by the share left, least first. A Redis read error answers 503 rather than zeros. `usage` needs a count for keys that have no limits, which no existing counter provides (limit counters exist only for configured limits; the request log is optional and holds raw tokens). The gateway now counts each key lineage's requests and weighted tokens per UTC day and month in Redis, at the points the limits engine counts them: a request once it passes the rate limits, its tokens in post-flight accounting. The writes fail open. The month's cost comes from gateway_logs, as the cost reports read it, and is null without ClickHouse. The endpoint sits behind a read-only variant of the API-key middleware: same lookup, checks and refusals, but it neither updates last_used_at (a poller must not keep an idle key past its inactivity timeout) nor arms the early-cancel request row. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The answer used to mix the key's own limits with its owner's, next to a
usage count for the key alone, so a client could not tell whose room the
numbers described. It now picks one subject and says which in a
top-level `scope`:
- "key" when the key has limits of its own on the AI gateway: `limits`
lists only those, counted for the key, and `usage` is the key's own.
- "user" when it has none: `limits` lists the owner's effective limits
(roles, including team-granted ones, merged most-restrictive, then
overrides), counted for everything the owner does, and `usage` is the
owner's total over all of their keys, `cost_usd_month` included. With
no limits on the owner either, `limits` is empty.
The limits still come from `limits_for_ai_gateway`, filtered to the
chosen subject, and `used` is read as before. A key whose only limit is
on the MCP gateway answers for its owner.
The owner's totals need a counter of their own: the user-subject
counters that exist are limit counters, kept only for configured limits
and only since they were configured, so they cannot give a day or month
total. The per-key counters become per-key and per-user ones in one
module (`limits::usage`, was `key_usage`): every counted request and its
tokens land on the key's and the owner's day and month hashes, all four
under the owner's `{user:<id>}` hash tag, one script per count, at the
same two points as before.
An answer from the response cache now counts its tokens on these
counters too, weighted as a call's are, so the usage agrees with the
limits once those count cache hits as well.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The backend has long held rate-limit rules and budgets on a key lineage
and enforced them, but the console could only set limits on users and
roles. The key's edit dialog now has a Limits tab: the key's rules
(requests or weighted tokens over 1m … 1w) and budgets (daily, weekly,
monthly weighted tokens), each with what it has used, read through the
existing `/api/admin/limits/api_key/{id}/rules|budgets|usage` endpoints
and added, changed and removed through the existing POST and DELETE
ones.
The tab reuses the user limits tab's parts rather than copying them:
the usage meter, expiry cell and rule label are exported from it, and
its override drawer becomes `LimitDrawer`, taking the subject to write
to. For a key the drawer defaults to no expiry and, when editing,
prefills the row's expiry and reason. In an edit the slot fields (type,
gateway, metric, window, period) are locked for both users and keys:
the row is stored by that slot, so changing one created a second row
instead of editing the first.
Permissions are the endpoints' own: the tab shows to holders of
`rate_limits:read` and disappears when the endpoints answer 403 (a
team-scoped grant never covers the caller's own keys); without
`rate_limits:write` it is read-only. Everyone else sees the dialog as
before.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Resolves the overlap with #94: cached answers are accounted once, through post_flight_account, which also records the key and user usage counters (the separate count_cache_hit_usage is gone); usage::record_request runs after check_limits in the new budget, access, limits order; current_spend propagates a counter read error. Removes the gateway RateLimiter that was built into GatewayState and never called. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Two things, both decided for this release:
GET /v1/usageon the gateway port tells a client holding a gateway API key what room it has left, about one subject, named in a top-levelscope.GET /v1/usage: the rule"scope": "key"when the key has any limit of its own on the AI gateway (a rate-limit rule or a budget on its lineage).limitslists only the key's limits, each withusedcounted for the key;usageis the key's own (across its rotations);cost_usd_monthis the key's."scope": "user"when the key has none.limitslists the owner's effective limits (roles, including team-granted roles, merged most-restrictive, then the user's overrides), each withusedcounted for everything the owner does;usageis the owner's total over all of their keys;cost_usd_monthis summed over all of their keys."scope": "user","limits": [], the owner's totals.scopeis always present and is exactly"key"or"user"; each limit entry keeps its ownscopetoo (always equal to the top-level one). A key whose only limit is on the MCP gateway answers for its owner.{ "scope": "key", "usage": {"requests_today": 12, "tokens_today": 48210, "requests_month": 340, "tokens_month": 1290455, "cost_usd_month": 3.82}, "limits": [ {"scope": "key", "kind": "tokens", "window": "daily", "window_secs": null, "limit": 100000, "used": 48210, "resets_at": "2026-10-12T00:00:00Z"}, {"scope": "key", "kind": "requests", "window": "1m", "window_secs": 60, "limit": 60, "used": 4, "resets_at": null} ], "expires_at": "2026-12-31T00:00:00Z" }A holder of a key without limits of its own sees the account's totals: the owner's requests, weighted tokens and cost today/this month across every key, and the owner's limits with what the whole account has used of them. A key handed to someone else should carry limits of its own if the owner's numbers are not theirs to see.
limitsis built withlimits_for_ai_gateway(what the pre-flight stages enforce) and then filtered to the chosen subject.usedis read from the same Redis counters the same way (sliding::read_count,budget::read_spend), without charging them. Ordered by the share left, least first; ties put shorter sliding windows first, then calendar periods, then requests before tokens. MCP-surface limits are not listed.usagecomes from per-key and per-user day/month counters (limits::usage, renamed from the earlierkey_usage). The user-subject counters that exist are limit counters: kept only for configured limits and only since they were configured, so they cannot give a day or month total for a user, and new per-user counters sit next to the per-key ones. Every counted request writes four hashes under the owner's hash tag in one script:usage:{user:<owner>}:api_key_lineage:<lineage>:daily:<date>/:monthly:<month>andusage:{user:<owner>}:user:<owner>:daily:<date>/:monthly:<month>, TTL 2× the period. They count at the same points as before: a request once it passes the rate limits, its weighted tokens inpost_flight_account. Answers from the response cache now count their tokens (the stored prompt + completion totals, weighted for the model as a call's are), so usage agrees with limits and budgets once those count cache hits too (see "Merging with fix/limit-enforcement"). Writes fail open withgateway_usage_count_fail_open_total.cost_usd_month:sum(cost_usd)fromgateway_logssince the 1st (UTC), byapi_key_lineage_idfor a key and byuser_idfor a user, as the cost reports sum it;nullwithout ClickHouse or on a failed read.expires_at, errors (same 401/403 as a model request, 503 when the counters can't be read) and the read-only middleware (nolast_used_atupdate, no early-cancel request row, nothing charged) are unchanged from the first commit.Key limits in the console
GET /api/admin/limits/api_key/{id}/rules|budgets|usage,POST …/rules|budgets,DELETE …/rules/{id}|budgets/{id}. No new endpoint, no backend permission change.LimitDrawer, which takes the subject (userorapi_key). For a key it defaults to no expiry and, in an edit, prefills the row's expiry and reason.rate_limits:read, and disappears if the endpoints answer 403 (e.g. a team-scoped grant, which never covers the caller's own keys). Withoutrate_limits:writeit is read-only (no add, edit or delete). Everyone else sees the dialog exactly as before.pnpm check:i18nparity holds). The Configuration Guide line for/v1/usagenow states the rule.Security-relevant
rate_limits:read|write+assert_scope_for_subject).Tests
handlers::key_usage): neither has limits →user,[]; key without limits (an MCP-only key rule included) → owner's limits only; key with limits while the owner has some → key's limits only; a key budget alone makes the answer the key's; ordering and window labels.limits::usage: the four hashes share one cluster slot; day/month turnover.limits.rs):scope: key, only the key's limits and usage, although the owner has limits (one with less left) and another key of the owner is used;scope: user, the owner's role / team-role / override limits withusedover both keys, andusagethe owner's totals including the other key's requests; both keys get the same answer;scope: user,limits: [], totals over two keys (viax-api-key);scope: key) and on the owner (scope: userfrom another key);last_used_at.analytics_clickhouse.rs):cost_usd_monthis the key's for a key with limits (2 calls) and the owner's for a key without (3 calls).key-limits-tab.test.tsx(lists rules and budgets with usage; empty state; read-only without write; add posts a permanent rule toapi_key/{id}; edit keeps the slot, expiry and reason; delete a budget after confirmation) andapi-keys.test.tsx(the tab in the edit dialog; no tab withoutrate_limits:read; no tab on 403).Not covered by a test: the 503 path (the harness has no Redis fault injection). The new tab has not been checked by eye in a browser.
Merging with fix/limit-enforcement
That branch makes cache hits call
post_flight_account, which (here) records the usage tokens. Both branches touch the cache-hit block inproxy/generate.rsand the pre-flight inproxy/pipeline.rs, so whichever lands second gets conflicts there. Resolve by:count_cache_hit_usagecall in the cache-hit block (and the now-unused function inaccounting.rs):post_flight_accountcounts those tokens, and keeping both would count them twice;usage::record_requestcall right aftercheck_limits, wherever that stage ends up in the new order (budget → access → limits).New messages
None (no core message codes involved).
🤖 Generated with Claude Code