Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
69 changes: 69 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -54,6 +54,74 @@ tokens. The legacy monthly token quota is gone.
tokens the cached answer records, weighted like any answer (all of its
input as plain input). The cost a cache hit reports stays 0. Callers that
lean on the cache reach their token limits and budgets sooner.
- **Each API key and each user get day and month usage counters in Redis.**
An AI-gateway request made with a key now also counts on the key's
`usage:{user:<owner id>}:api_key_lineage:<lineage id>:daily:<date>` and
`…:monthly:<month>` and on its owner's
`usage:{user:<owner id>}:user:<owner id>:daily:<date>` and
`…:monthly:<month>`, hashes that expire two periods after their last write.
That is one more Redis script per request once it passes the rate limits,
and one when its tokens are counted, after the call or when it is answered
from the response cache. The counters start empty, so the `usage` that
`GET /v1/usage` reports counts from the upgrade on. No setting, database
schema or Helm value changes.

### Added

- **`GET /v1/usage` tells a client what room its API key has left.** Called
on the gateway port with a gateway API key, as a model request is, it
answers about one of two subjects, named in `scope`:

- `"key"` when the key has limits of its own on the AI gateway: `limits`
lists only the key's limits, each with `used` counted for the key, and
`usage` is the key's own (across its rotations).
- `"user"` when it has none: `limits` lists its owner's effective limits
(roles, including those a team grants, merged most-restrictive, then
the user's overrides), each with `used` counted for everything the
owner does, and `usage` is the owner's total over all of their keys.
With no limits on the owner either, `limits` is empty.

```json
{
"scope": "key",
"usage": {"requests_today": 12, "tokens_today": 48210,
"requests_month": 340, "tokens_month": 1290455,
"cost_usd_month": 3.82},
"limits": [
{"scope": "key", "kind": "tokens", "window": "daily", "window_secs": null,
"limit": 100000, "used": 48210, "resets_at": "2026-10-12T00:00:00Z"},
{"scope": "key", "kind": "requests", "window": "1m", "window_secs": 60,
"limit": 60, "used": 4, "resets_at": null}
],
"expires_at": "2026-12-31T00:00:00Z"
}
```

`usage` holds the requests the rate limits let through and the weighted
tokens limits count, answers from the response cache included, for the
UTC day and month, and the cost this month from the request log (`null`
without ClickHouse). Every limit's `used` is read from the counter that
refuses requests, and the one with the least left comes first. `window`
is a rate limit's sliding window (`1m`, `5m`, `1h`, `5h`, `1d`, `1w`, with
its length in `window_secs`) or a budget's calendar period (`daily`,
`weekly`, `monthly`, with its end in `resets_at`, UTC). Limits on the MCP
gateway are not listed. `expires_at` is the key's expiry or the end of
its rotation grace period, whichever comes first. A key a model request
would refuse gets the same `401` or `403`, and `503` means the counters
could not be read. Calling it charges no limit, writes no request log row
and is not a use of the key: `last_used_at` stays as it was, so polling
does not keep an idle key from its inactivity timeout. Whoever holds a key
without limits of its own sees its owner's totals and limits; a key handed
to someone else should carry limits of its own. The console's
Configuration Guide lists the endpoint with the other gateway endpoints.
- **An API key's own limits are edited in the console.** The key's edit
dialog has a Limits tab: rate limits (requests or weighted tokens over
1m, 5m, 1h, 5h, 1d or 1w) and budgets (daily, weekly or monthly weighted
tokens), each with what it has used, added, changed and removed through
the existing limits endpoints. They follow the key across rotations.
Reading needs `rate_limits:read` and changing `rate_limits:write`, in a
scope that covers the key, as before; without write access the tab is
read-only, and without read access the dialog is unchanged.

### Fixed

Expand Down Expand Up @@ -83,6 +151,7 @@ tokens. The legacy monthly token quota is gone.
- The `security.rate_limit_fail_closed` hint in Settings names everything
the setting now covers.


## [3.4.0] — 2026-10-11

Conversations with reasoning models behind Chat-format upstreams keep their
Expand Down
56 changes: 42 additions & 14 deletions crates/common/src/limits/budget.rs
Original file line number Diff line number Diff line change
Expand Up @@ -190,24 +190,52 @@ pub async fn current_spend(
let now = Utc::now();
let mut out = Vec::with_capacity(caps.len());
for cap in caps {
let key = build_key(
cap.subject_kind.as_str(),
cap.subject_id,
cap.period.as_str(),
now,
);
let v: Option<i64> = redis.get(&key).await?;
out.push(CapStatus {
cap_id: cap.id,
subject_kind: cap.subject_kind,
subject_id: cap.subject_id,
current: v.unwrap_or(0),
limit: cap.limit_tokens,
});
out.push(cap_status(cap, spend(redis, cap, now).await?));
}
Ok(out)
}

/// [`current_spend`] at `now`, except that a Redis error is returned
/// rather than read as nothing spent.
pub async fn read_spend(
redis: &Client,
caps: &[BudgetCap],
now: DateTime<Utc>,
) -> Result<Vec<CapStatus>, fred::error::Error> {
let mut out = Vec::with_capacity(caps.len());
for cap in caps {
out.push(cap_status(cap, spend(redis, cap, now).await?));
}
Ok(out)
}

/// The counter of `cap`'s period containing `now`; a period nothing has
/// been spent in yet has no counter.
async fn spend(
redis: &Client,
cap: &BudgetCap,
now: DateTime<Utc>,
) -> Result<i64, fred::error::Error> {
let key = build_key(
cap.subject_kind.as_str(),
cap.subject_id,
cap.period.as_str(),
now,
);
let v: Option<i64> = redis.get(&key).await?;
Ok(v.unwrap_or(0))
}

fn cap_status(cap: &BudgetCap, current: i64) -> CapStatus {
CapStatus {
cap_id: cap.id,
subject_kind: cap.subject_kind,
subject_id: cap.subject_id,
current,
limit: cap.limit_tokens,
}
}

/// Add `weighted_tokens` to every cap counter in the slice and
/// return the new running totals so the caller can fire alerts on
/// crossings. Soft cap: this never blocks the request — it only
Expand Down
6 changes: 6 additions & 0 deletions crates/common/src/limits/mod.rs
Original file line number Diff line number Diff line change
Expand Up @@ -61,6 +61,7 @@ use uuid::Uuid;

pub mod budget;
pub mod sliding;
pub mod usage;
pub mod weight;

// ----------------------------------------------------------------------------
Expand Down Expand Up @@ -1162,6 +1163,10 @@ pub struct RequestLimits {
/// The user the request runs as. Every rate-limit counter below
/// carries their Redis Cluster hash tag (`sliding::counter_key`).
pub owner: Uuid,
/// The lineage of the API key the request came with, whether or not
/// it has limits of its own: its `usage` counters, and its owner's,
/// count the request.
pub key_lineage: Option<Uuid>,
pub rules: Vec<RateLimitRule>,
pub caps: Vec<BudgetCap>,
}
Expand All @@ -1181,6 +1186,7 @@ impl RequestLimits {
) -> Self {
let mut out = Self {
owner,
key_lineage: key.map(|(lineage, _)| lineage),
..Self::default()
};
out.add(surface, RateLimitSubject::User, owner, user);
Expand Down
18 changes: 15 additions & 3 deletions crates/common/src/limits/sliding.rs
Original file line number Diff line number Diff line change
Expand Up @@ -503,16 +503,28 @@ pub async fn record_at(
/// Returns 0 on Redis error so the UI can fall back to "no data"
/// rather than 500. Real failures are logged.
pub async fn current_count(redis: &Client, rule: &ResolvedRule) -> i64 {
use fred::interfaces::HashesInterface;
match redis.hgetall::<HashMap<String, i64>, _>(&rule.key).await {
Ok(buckets) => window_sum(&buckets, chrono::Utc::now().timestamp(), rule.bucket_secs),
match read_count(redis, rule, chrono::Utc::now().timestamp()).await {
Ok(count) => count,
Err(e) => {
tracing::warn!(key = %rule.key, "rate-limit usage read failed: {e}");
0
}
}
}

/// What the window that ends at `now_secs` holds — the sum [`admit`]
/// compares with the limit — without changing it. A Redis error is
/// returned, not read as an empty window.
pub async fn read_count(
redis: &Client,
rule: &ResolvedRule,
now_secs: i64,
) -> Result<i64, fred::error::Error> {
use fred::interfaces::HashesInterface;
let buckets = redis.hgetall::<HashMap<String, i64>, _>(&rule.key).await?;
Ok(window_sum(&buckets, now_secs, rule.bucket_secs))
}

/// The sum of the buckets inside the window that ends at `now_secs`.
fn window_sum(buckets: &HashMap<String, i64>, now_secs: i64, bucket_secs: i32) -> i64 {
let bucket_secs = i64::from(bucket_secs);
Expand Down
Loading
Loading