Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
17 changes: 17 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -114,6 +114,23 @@ All notable changes to this project are documented in this file. The format is b
module defers past the request (#364). The `tenants` module resolves a
tenant from the subdomain (`subdomain_base`), anonymous visitors included
(#363).
- **Request-path performance** (Postgres load test,
[docs/perf/2026-10-08-postgres-loadtest.md](docs/perf/2026-10-08-postgres-loadtest.md)).
Route matching skips any included router whose routes cannot match the path:
FastAPI ≥ 0.140 regex-tested nearly all ~190 routes per request, ~30% of a
cheap request's CPU (`/health` 1.98 → 1.46 ms). The guards are built at
startup, which also takes FastAPI's lazy per-route build off the first
request after boot (360 → 32 ms). GZip runs at level 5 instead of 9: about
1% larger output for 2.4× less CPU. Trivial sync FastAPI dependencies
(`get_permission_registry`, `get_feature_flag_registry`, the `file_storage`,
`tenants` and `users` accessors) are now `async`, so they no longer
take a threadpool round-trip. `/admin/users` counts its status cards in one
query instead of three.
- `SetupMiddleware` refreshes its cached verdict single-flight. When the 5 s
TTL lapsed under load, every in-flight request ran the setup steps itself,
each checking out a pooled connection. Behind a saturated pool those
checkouts queued and fed `QueuePool limit … reached` timeouts. The expired
*complete* verdict now keeps answering while one refresh runs.

### Security
- The tenant header (`tenant_header`) is no longer honoured for an
Expand Down
2 changes: 2 additions & 0 deletions docs/framework/middleware.md
Original file line number Diff line number Diff line change
Expand Up @@ -138,6 +138,8 @@ Emits a structured log line per request with method, path, status, duration, and

Starlette's, compressing any response body over `COMPRESSION_MIN_BYTES` (500). Placed inside `CorrelationId` and `RequestLogging` — which set headers and read request state — but outside everything that produces a body, **including the `/static` mount**, which is where it earns its place: the built CSS is ~139 KB raw against ~21 KB gzipped, and the JS bundle compresses about 3×. Uncompressed assets dominated cold page load, several times larger than anything on the server request path.

It runs at `COMPRESSION_LEVEL` 5, not Starlette's default of 9. Compression runs on the event loop that serves every other request, and level 9 buys almost nothing on JSON and HTML: a 76 KB first-load page comes out 1.2% smaller than at level 5 for 2.4× the CPU (6.4 ms vs 2.7 ms). Pre-compressed static assets are served as-is and never re-compressed, so they keep whatever level they were built with.

### `SecurityHeadersMiddleware`

Sets conservative defaults: `X-Content-Type-Options: nosniff`, `Referrer-Policy: strict-origin-when-cross-origin`, `X-Frame-Options: SAMEORIGIN`, `X-XSS-Protection: 0` (the legacy auditor is disabled in favour of CSP), plus a default CSP and — outside development — HSTS. In development the CSP is widened for the Vite dev origin and HSTS is suppressed. Modules that load assets from an external origin extend the policy through the [`register_csp_sources`](lifecycle.md#register_csp_sourcesregistry) hook; both the dev and production variants honor those origins. Override on a per-route basis with your own response headers.
Expand Down
109 changes: 109 additions & 0 deletions docs/perf/2026-10-08-postgres-loadtest.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,109 @@
# Backend load test on Postgres

**Date:** 2026-10-08
**Harness:** [tests/loadtest/](../../tests/loadtest/README.md) (locust + faker seed)

## Environment

| | |
|---|---|
| Machine | linux, 24 cores, shared `dev-services` Postgres (Docker, `max_connections=100`) |
| Database | Postgres, fresh DB, 10 000 users (9 000 role assignments), 100 000 audit rows, `ANALYZE`d |
| Modules | `Auth, FeatureFlags, Settings, FileStorage, Users, AuditLog, BackgroundTasks, Dashboard, Permissions, Tenants, SiteLock, Branding` |
| Server | `uvicorn host.main:app`, development mode, one worker unless stated |
| Load | `locustfile.py` `AuthedUser` mix, 300 users, spawn 100/s, 60 s |
| Micro | in-process `httpx.ASGITransport`, sequential, 200 warm requests per URL, median |

## Saturated throughput (300 users, 1 worker)

| | req/s | p50 | p95 | failures |
|---|---:|---:|---:|---:|
| Before | 117.7 | 330 ms | 9.7 s | 1 (`QueuePool limit … reached`) |
| After | 123.5 | 330 ms | 9.5 s | 3 |
| After, `SM_DB_POOL_PRE_PING=false` | 127.1 | 310 ms | 9.8 s | 1 |
| After, 4 workers, pool 5 + 10 | **520.9** | **180 ms** | **790 ms** | **0** |

One worker is **CPU-bound**: the process sat at 98% of one core while
Postgres used about one core. The p95 of ~10 s is queueing, not slow
queries. Weighting each endpoint's measured CPU by the locust mix gives
~8.7 ms per request, a ceiling of ~115–125 req/s on one core, which matches what
was observed. Behind a saturated event loop each request holds its pooled
connection far longer than its queries take, so the pool (10 + 20) runs dry
and requests wait up to `pool_timeout` (30 s).

So throughput is a deployment lever, as in the June campaign: four workers
with the pool cut so that `workers × (pool_size + max_overflow)` stays under
`max_connections` gave 4.4× the throughput and a 12× better p95 with no
errors. See `docs/reference/deployment.md`.

## Per-request cost (in-process, ms, median)

Wall time; the DB-backed rows include Postgres round-trips and are noisy on the
shared instance.

| URL | before | after |
|---|---:|---:|
| `/health` | 1.98 | 1.46 |
| `/api/feature_flags/` | 2.19 | 1.90 |
| `/api/permissions/` | 3.00 | 2.52 |
| `/api/dashboard/stats` | 2.53 | 2.12 |
| `/dashboard/` (Inertia) | 3.67 | 3.24 |
| `/admin/users/` (Inertia) | 28.9 | 25.3 |
| `/api/users/admin` | 11.2 | 11.0 |
| `/api/audit_log/` | 22.0 | 24.1 |
| `/api/settings/modules` | 8.58 | 8.13 |

Main-thread CPU for the heavy endpoints: `/admin/users/` 17.5 ms,
`/api/users/admin` 10.5 ms, `/api/audit_log/` 8.4 ms, `/api/settings/modules`
8.1 ms. These make up most of the mix's CPU.

## Fixed

1. **Route matching was ~30% of a cheap request's CPU.** FastAPI ≥ 0.140
keeps every `include_router` as a lazy `_IncludedRouter`. The app root
holds ~35 of them (one API and one view router per module), and none
filters on its prefix, so a request regex-tested nearly all ~190 routes. It
also re-walked each subtree's route version and ran `_match` twice on the
router that matched. `simple_module_hosting._route_guard` gives each
top-level include a prefix guard derived from its effective route paths and
re-derived when FastAPI's route version changes. Saves ~0.5 ms per request.
2. **Cold start.** The first request to reach a router builds every one of its
routes' dependants (signatures and pydantic adapters). The guards are built
at lifespan start, which does that work during boot: the first request
dropped from 360 ms to 32 ms, and startup grew by about the same amount.
3. **GZip at level 9.** Starlette's default. A 76 KB first-load page cost
6.4 ms of event-loop CPU at level 9 against 2.7 ms at level 5, for 1.2%
less output. Now level 5.
4. **Setup-gate stampede.** The verdict TTL (5 s) lapsing under load made
every in-flight request run `has_administrator` on its own pooled
connection, which fed the pool exhaustion. Now single-flight, and an
expired *complete* verdict answers while the refresh runs.
5. **Threadpool hops.** Trivial sync dependency getters made FastAPI dispatch
each one to the threadpool. They are now `async def`.
6. **`/admin/users/` status cards** issued three `COUNT(*)` round-trips; now
one `COUNT(*) FILTER (WHERE …)` scan.

## Not changed: findings and follow-ups

- **`pool_pre_ping`** runs `BEGIN; ; ROLLBACK` (three round-trips) on every
checkout. That is ~1.4 ms of held-connection time per DB request against a
331 µs bare round-trip, plus ~4.5% of request CPU. Turning it off gave +3%
throughput. It stays on by default, since it is what lets a pool survive a
database restart, but a host with a stable database and its own retry layer
can set `SM_DB_POOL_PRE_PING=false`.
- **Soft-delete/tenant statement filter** (`query_filter.filter_statements`)
attaches `with_loader_criteria` for every soft-delete model to every SELECT
and walks subqueries: ~0.35 ms of CPU per statement. Compiled-cache hits
are unaffected (every statement measured was a `CACHE_HIT`). It is the
isolation boundary from #332, so it needs a design-level review rather than
a perf patch.
- **Audit-log `COUNT(*)`** over 100k rows takes ~15 ms on this instance and
dominates `/api/audit_log/`. An estimated or capped count would be a UX
decision.
- **`/api/settings/modules`** spends ~8 ms rebuilding the view of ~110
settings fields per call. It is a low-traffic admin screen that the mix
over-weights, and its output depends on live env, DB overrides and secret
masking, so it is left uncached.
- Inertia payloads: the first navigation per session carries the full i18n
catalog (~75 KB of a 76 KB page); later navigations send 4.3 KB. A client
without a cookie jar (curl) sees 76 KB every time, so measure with one.
4 changes: 4 additions & 0 deletions framework/hosting/simple_module_hosting/_lifespan.py
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,7 @@

from fastapi import FastAPI

from simple_module_hosting._route_guard import guard_included_routers
from simple_module_hosting.migrations import migration_status
from simple_module_hosting.setup_gate import STEP_MIGRATIONS

Expand Down Expand Up @@ -171,6 +172,9 @@ async def lifespan(app: FastAPI) -> AsyncGenerator[None, None]:
exc_info=True,
)
app.state.deferred_startup.append(mod)
# Here rather than in create_app: the host includes its own routers
# after create_app returns, and those should be guarded too.
guard_included_routers(app)
yield
for mod in reversed(modules):
await mod.on_shutdown(app)
Expand Down
9 changes: 8 additions & 1 deletion framework/hosting/simple_module_hosting/_phase_helpers.py
Original file line number Diff line number Diff line change
Expand Up @@ -58,6 +58,11 @@
# Below this, gzip framing costs more than it saves. Starlette's own default.
COMPRESSION_MIN_BYTES = 500

# Starlette defaults to 9, which buys almost nothing on JSON and HTML: a 76 KB
# first-load page compresses 1.2% smaller at 9 than at 5, for 2.4x the CPU
# (6.4 ms vs 2.7 ms), spent on the event loop that serves every other request.
COMPRESSION_LEVEL = 5

# Re-exported for back-compat: static-file serving now lives in static_files.
ImmutableStaticFiles = PrecompressedStaticFiles

Expand Down Expand Up @@ -176,7 +181,9 @@ def install_middleware(
# matters most: the built CSS is ~139 KB raw and ~21 KB gzipped, and the
# JS bundle compresses about 3x. Uncompressed assets dominated cold page
# load, several times larger than anything on the server request path.
app.add_middleware(GZipMiddleware, minimum_size=COMPRESSION_MIN_BYTES)
app.add_middleware(
GZipMiddleware, minimum_size=COMPRESSION_MIN_BYTES, compresslevel=COMPRESSION_LEVEL
)
# Right inside RequestLogging (so a 413 is logged) and outside GZip and the
# whole module tier: an oversized body is refused before anything reads it.
app.add_middleware(
Expand Down
98 changes: 98 additions & 0 deletions framework/hosting/simple_module_hosting/_route_guard.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,98 @@
"""Skip whole included routers whose routes cannot match the request path.

FastAPI >= 0.140 keeps every ``include_router`` as a lazy ``_IncludedRouter``
placeholder instead of copying its routes onto the parent. Matching a request
then asks each placeholder in turn, and a placeholder that cannot match still
walks its subtree's route version and regex-tests every route in it. With one
API and one view router per module the app root holds ~35 placeholders, so a
request for a module near the end of the list tested nearly all ~190 routes:
route matching was ~30% of the CPU of a cheap authenticated request.

Every route under a placeholder shares the common prefix of its effective
paths. A request whose path lacks that prefix cannot match any of them —
neither fully nor partially — so the guard answers ``Match.NONE`` without
descending. Anything else is delegated unchanged, so a match still takes
FastAPI's own path.

The prefix is derived from the routes, not from ``APIRouter.prefix``: a route
added with ``add_route`` does not carry its router's prefix. It is recomputed
whenever FastAPI's route version for that router changes, so a router mutated
after the app was built is never filtered against a stale prefix.

Relies on ``fastapi.routing._IncludedRouter``; on a FastAPI without it (which
flattened routes eagerly and needs no guard) this is a no-op.
"""

from __future__ import annotations

import os
from typing import Any

from fastapi import FastAPI
from starlette._utils import get_route_path
from starlette.routing import Match
from starlette.types import Scope

try: # private, and absent before FastAPI 0.140
from fastapi.routing import _IncludedRouter
except ImportError: # pragma: no cover - older FastAPI flattens eagerly
_IncludedRouter = None # type: ignore[assignment,misc]

__all__ = ["guard_included_routers"]


def _effective_paths(included: Any) -> list[str]:
paths: list[str] = []
for ctx in included.effective_route_contexts():
route = ctx.starlette_route
paths.append((getattr(route, "path", "") if route is not None else ctx.path) or "")
return paths


def _common_prefix(paths: list[str]) -> str:
"""The literal leading part every path shares — never past a ``{param}``."""
if not paths:
return ""
return os.path.commonprefix(paths).split("{", 1)[0]


class _PrefixGuard:
"""Replacement ``matches`` for one ``_IncludedRouter`` instance."""

def __init__(self, included: Any) -> None:
self._included = included
self._matches = included.matches
self._version: int | None = None
self._prefix = ""

def _current_prefix(self) -> str:
version = self._included.original_router._get_routes_version()
if version != self._version:
self._prefix = _common_prefix(_effective_paths(self._included))
self._version = version
return self._prefix

def __call__(self, scope: Scope) -> tuple[Match, Scope]:
prefix = self._current_prefix()
# "/" (or "") is shared by every path, so it would filter nothing.
if len(prefix) > 1 and not get_route_path(scope).startswith(prefix):
return Match.NONE, {}
return self._matches(scope)


def guard_included_routers(app: FastAPI) -> None:
"""Install a prefix guard on each router included directly into ``app``.

Computing a guard's prefix resolves its router's effective routes, which is
also what FastAPI does lazily on the first request to reach them — building
every route's dependant, signature and pydantic adapters. Doing it here
moves that one-time cost (hundreds of ms for a full module set) from the
first request after boot into startup.
"""
if _IncludedRouter is None:
return
for route in app.router.routes:
if isinstance(route, _IncludedRouter) and not isinstance(route.matches, _PrefixGuard):
guard = _PrefixGuard(route)
guard._current_prefix()
route.matches = guard # type: ignore[method-assign]
34 changes: 30 additions & 4 deletions framework/hosting/simple_module_hosting/setup_gate.py
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,7 @@

from __future__ import annotations

import asyncio
import logging
import time

Expand Down Expand Up @@ -128,6 +129,7 @@ def __init__(self, app: ASGIApp) -> None:
self.app = app
self._verdict: bool | None = None
self._verdict_expires: float = 0.0
self._refresh: asyncio.Future[bool] | None = None
self._announced = False

async def _is_complete(self, registry, starlette_app) -> bool:
Expand All @@ -149,15 +151,39 @@ async def _is_complete(self, registry, starlette_app) -> bool:
The cache lives on the middleware instance rather than on ``app.state``
so it cannot leak between two apps built in the same process — the test
suite builds many — and is discarded with the app that owns it.

Refreshes are single-flight. When the TTL lapses under load, every
request in flight used to run its own evaluation, each checking out a
pooled connection; behind a saturated pool those checkouts queued for
seconds and starved the requests that actually needed the database.
Now one evaluation runs, and while it does, the expired *complete*
verdict keeps answering — the refresh still lands every TTL, so a lost
administrator still brings the wizard back. With no complete verdict to
fall back on, callers share the in-flight evaluation instead.
"""
now = time.monotonic()
if self._verdict and now < self._verdict_expires:
return True
verdict = await registry.is_setup_complete(starlette_app)
if verdict:
refresh = self._refresh
if refresh is None:
refresh = asyncio.ensure_future(registry.is_setup_complete(starlette_app))
refresh.add_done_callback(self._refresh_done)
self._refresh = refresh
elif self._verdict:
return True
# Shielded so a caller that is cancelled (client gone) does not cancel
# the evaluation every other waiter is sharing.
return await asyncio.shield(refresh)

def _refresh_done(self, refresh: asyncio.Future) -> None:
self._refresh = None
if refresh.cancelled() or refresh.exception() is not None:
return
if refresh.result():
self._verdict = True
self._verdict_expires = now + _VERDICT_TTL_SECONDS
return verdict
self._verdict_expires = time.monotonic() + _VERDICT_TTL_SECONDS
else:
self._verdict = False

async def __call__(self, scope: Scope, receive: Receive, send: Send) -> None:
if scope["type"] != "http":
Expand Down
Loading
Loading