Skip to content

Repository files navigation

One the Gateway

English · 한국어

A self-hosted LLM gateway for teams. Keep Provider credentials on the server and give clients Gateway tokens and model aliases.

Built with Rust and MySQL 8.4. The web console is embedded in the binary; serving it needs no frontend build, Node.js or Python runtime.

What it does

  • Accepts OpenAI Chat Completions, OpenAI Responses and Anthropic Messages. Translates supported text, function calls/results and incremental SSE between configured Provider protocols.
  • Connects OpenAI, Anthropic, OpenCode and custom OpenAI-compatible servers. Project-supported Codex OAuth connects an account in the administrator console to use its available Codex allowance; it is not OpenAI endorsement or live-account certification.
  • Imports original Provider model metadata and creates new aliases with user access by default. Existing administrator policies and manually configured capabilities are preserved.
  • Applies user/token scopes, role-based routing and atomic UTC-day request quotas. Stores encrypted Provider credentials and hashed Gateway tokens.
  • Supports GitHub Organization/team login and enterprise OIDC. SAML requires an external OIDC broker; there is no native SAML parser.
  • Discovers Provider/SSO Rust adapter files automatically at build time. Add an adapter without editing core registries or the console.
  • Pools up to 32 isolated keys/accounts per Provider with shared models/policy and remaining-capacity-first selection. Pauses accounts independently; never replays dispatched inference.
  • Monitors per-account UTC-day requests/tokens and available upstream quotas: Codex/OpenCode Go usage endpoints and observed OpenAI/Anthropic rate-limit headers. Unknown or unsupported balances stay unknown.
  • Provides an English/Korean console for Providers, routing, tokens, users, SSO, monitoring and request records.
  • Offers opt-in configuration installers for Pi, Codex, Claude Code and Hermes without replacing their existing logins or default models.

Version 0.1.0. This is a limited protocol implementation, not a transparent proxy. Multimodal input, signed/encrypted reasoning, strict structured output and arbitrary upstream request fields are unsupported. There is no automatic retry/failover after inference dispatch or response cache. See the compatibility matrix before connecting a client. Local test coverage is not live Provider, agent or enterprise IdP certification.

Quick start

1. Prerequisites

  • Current stable Rust toolchain and MySQL 8.4.
  • An empty database and a dedicated database account with schema CREATE/ALTER/INDEX privileges, in addition to application read/write access.
  • A sign-in provider: GitHub Organization/team + OAuth App, enterprise OIDC, or a compiled SSO adapter supporting initial setup. For GitHub, use <BASE_URL>/auth/github/callback and permit read:org; for OIDC, use <BASE_URL>/auth/sso/<slug>/callback and configure verified administrator role/group claims.

Initial setup uses a web wizard; GitHub is not required when another supported SSO provider is selected. For a disposable local database and existing-database migration guidance, see operations.

git clone https://github.com/2tle/one-gateway.git
cd one-gateway
cargo build --locked --release

2. Retain configuration

Generate an encryption key once with openssl rand -base64 32 and retain it securely. Put trusted shell assignments in a private .env at the repository root:

DATABASE_URL='mysql://gateway:YOUR_DATABASE_PASSWORD@127.0.0.1/gateway'
BASE_URL='http://localhost:3000'
GATEWAY_ENCRYPTION_KEY='YOUR_RETAINED_BASE64_KEY'
chmod 600 .env
bash scripts/run-gateway.sh

Never commit .env, print secrets in logs or enable shell tracing. The wrapper loads the same configuration on every invocation; the Rust binary itself does not load .env. Use GATEWAY_ENV_FILE=/absolute/path/gateway.env for an external trusted file, or inject the environment through a secret store. A present file overrides inherited values. Because the file is sourced as shell code, never use an untrusted file.

Keep the original encryption key across restarts, upgrades and replicas. Losing or changing it makes stored credentials unreadable. The wrapper does not install a boot service. Configure your service manager separately. Remote deployments require an HTTPS BASE_URL, a trusted TLS reverse proxy and explicit listen/proxy configuration; only loopback HTTP is accepted for development. See the deployment checklist.

3. Bootstrap and sign in

  1. Open <BASE_URL>/ (locally, http://localhost:3000/). An empty installation opens the setup wizard.
  2. Enter the one-time setup token printed in the server startup log. This opens a setup-only session, not an administrator account.
  3. Choose GitHub, OIDC or a registered custom SSO adapter. Enter its credentials, callback configuration and administrator policy. For OIDC, explicitly choose all subjects of a dedicated issuer or a subject allowlist; shared issuers should use an allowlist.
  4. Add one or more model Providers with their protocol, base URL, credentials and settings. Custom adapters appear automatically. OAuth accounts may be registered unconnected.
  5. Review and finish. SSO and Providers are saved together; the setup token/session are invalidated. Sign in through the configured SSO with an account matching its administrator policy.

Setup sessions last 30 minutes. Unsaved inputs stay only in page memory; refreshing loses them. If the session is lost or expires before completion, restart the server to obtain a new key. Initialize through one server instance; setup sessions are process-local. Completed/populated installations never reopen setup. No first-login administrator promotion occurs. The legacy API on localhost port 3001 remains compatible; never expose or proxy that port. Security, recovery and legacy setup.

4. Configure a model and issue a token

  1. Use the Providers registered during setup, or add more in the console. Registration does not verify inference or create routing aliases. For Codex OAuth, leave the credential empty, open Models → Start authorization, complete the device login and confirm it before syncing models. See account connection and recovery.
  2. Open its model catalog and sync models, or add a model manually. For a manual model, create an enabled alias and add the model as an enabled routing candidate. Catalog import is not an inference health check.
  3. Check the alias's routing candidates, minimum role, client protocols and capabilities. Agent use requires verified text, tools and streaming; missing Provider declarations default to text only.
  4. Issue a Gateway token. Save it immediately; it is shown only once.

The top-bar language selector switches between English and 한국어. Browser preference supplies the default, and an explicit selection is stored locally. API values, model IDs and original Provider metadata are not translated.

Add accounts and monitor usage

In Providers → Accounts / keys, add independently encrypted keys or connect additional Codex OAuth accounts. They share the root Provider's models and access policy. Usage separates recorded Gateway requests/tokens from upstream windows, resets, observation age and refresh failures. Positive observations expire for selection after 120 seconds; observed exhaustion blocks until reset. Multiple keys may share one upstream quota, so adding keys does not multiply it. See account APIs, quota sources and operational limits.

Call the Gateway

Use the Gateway token, not a Provider credential, and a configured alias as model:

export GATEWAY_TOKEN='gw_YOUR_TOKEN'
curl -fsS http://localhost:3000/v1/chat/completions \
  -H "Authorization: Bearer $GATEWAY_TOKEN" \
  -H 'Content-Type: application/json' \
  -d '{"model":"coding/default","messages":[{"role":"system","content":"You are a helpful assistant."},{"role":"user","content":"Hello"}],"max_completion_tokens":128,"stream":true}'
Endpoint Purpose
GET /v1/models Discover aliases accessible to this token; no inference quota charge
POST /v1/chat/completions OpenAI Chat Completions
POST /v1/responses OpenAI Responses
POST /v1/messages Anthropic Messages

All inference endpoints enforce the same authorization and quota policy. Quota is charged before Provider inference and is not refunded on later failure. Unsupported semantics fail explicitly instead of being silently removed. Read the API reference for request contracts.

Connect an agent

The console generates an opt-in setup command for an already installed CLI. The installer verifies the selected alias with a tool call and streamed follow-up before writing configuration. These two inference requests consume quota and may incur Provider charges.

export ONE_THE_GATEWAY_API_KEY='gw_YOUR_TOKEN'
export ONE_THE_GATEWAY_MODEL='coding/default' # Optional: choose an accessible alias.
(set -o pipefail; curl -fsS http://localhost:3000/scripts/install-pi.sh | bash)

Replace pi with codex, claude or hermes for another agent. Bash, Python 3 and the selected CLI must be installed on the client. Inspect scripts before running them and protect shell history/clipboard containing keys. Use the printed *-gateway launcher afterward; ordinary CLI commands retain their existing defaults. See installation and recovery for details and limits.

Operating limits

Environment variable Default Purpose
LISTEN_ADDR 127.0.0.1:3000 Application bind address
GATEWAY_MAX_IN_FLIGHT 300 Process-local concurrent inference limit
DATABASE_MAX_CONNECTIONS 32 MySQL pool size per process
DATABASE_ACQUIRE_TIMEOUT_SECONDS 5 Database pool acquisition timeout
GATEWAY_SHUTDOWN_GRACE_SECONDS 60 Shared HTTP/audit shutdown deadline

Inference admission saturation returns 503 before quota charging or Provider inference. SIGTERM/Ctrl-C closes admission and drains accepted work within the grace deadline. Expiry can interrupt unfinished requests/audits. These limits are not distributed Provider rate limits or a promise of 300 sustained real-model requests. Tune them with real deployment measurements; disable proxy SSE buffering and allow long-running responses. Full operational behavior and ranges.

Development and tests

cargo fmt --check
cargo clippy --locked --all-targets -- -D warnings
cargo test --locked
python3 tests/installers.py
python3 tests/startup.py
python3 tests/extensions.py
node tests/i18n.cjs

For the full database suite, use a disposable MySQL 8.4 server, never production:

TEST_DATABASE_URL=mysql://root:gateway-test@127.0.0.1:13306/mysql \
  cargo test --locked -- --include-ignored

CI runs Rust/MySQL, installer, startup, translation, browser/accessibility and Docker/Helm checks. Browser checks use mocked APIs; Node/Playwright/Axe are development-only tools. See test setup and browser commands.

Documentation

Contributing

See CONTRIBUTING.md for focused pull requests, local checks and safe bug reports. English and 한국어 contributions are welcome. Use synthetic credentials and disposable databases; never include real keys, tokens, request bodies or personal data. Contributions are licensed under Apache-2.0.

License

Apache License 2.0. Bundled third-party assets retain their own licenses: DOMPurify and IBM Plex Sans. Dependency licenses still apply when distributing binaries.

About

Lightweight, open-source enterprise LLM gateway in Rust — built for SSO, scalable routing, and pluggable LLM providers.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages