Configuration
This page explains how a running gateway finds its configuration: which files it reads, where those files live, and which layer wins when more than one names the same file. It is the conceptual companion to Adding a New Model (the field-by-field model registry reference) and Routing (what the routing engine does with the result).
For distribution manifests, branding, UI modules and backend extensions, see Distribution customization.
What the gateway reads at startup
Kind |
Holds |
|---|---|
|
The model registry: every model id the gateway serves and the upstream routes behind it |
|
Optional deployment-wide settings: health probing and a local/remote weight split |
|
Alert rules and thresholds |
Environment |
Everything secret or host-specific: credentials, database connection, feature switches |
Only the model registry is required to serve traffic. Without it the gateway
starts, serves /health, and answers GET /v1/models with an empty list.
Without a routing file each route keeps the weight the registry gives it, and
without an alerts file the built-in thresholds apply.
The files are not the whole story once a database is configured. The admin console writes providers, keys, routes, weights and per-model overrides to the operational store, and the gateway re-applies that state on top of the loaded registry at every start. Runtime configuration from the admin console covers that layer.
Where configuration lives
config/examples/ contains reference gateway configuration, including
models.openrouter.yaml and routing.minimal.yaml. A deployment keeps its
model registry, routing rules and alerts in its own configuration directory,
such as distributions/<name>/config/, and selects them through explicit
environment variables or an active distribution manifest.
Distribution customization
explains the manifest and overlay layout.
A provider, key, route or whole model can also be added from the admin console while the gateway runs. That state lives in Postgres and is described below.
How a gateway finds its config
Precedence is environment, then manifest, then built-in default.
Each kind of file is looked up on its own, in this order:
An explicit environment variable.
MODELS_CONFIG_PATH,ROUTING_CONFIG_PATH,ALERTS_CONFIG_PATH. (The older namesMODELS_CONFIGandROUTING_CONFIGare still accepted; the canonical*_CONFIG_PATHname wins when both are set.)The distribution manifest’s
paths:section — only whenDISTRIBUTION_CONFIG_MODE=active. By default a manifest is only checked, not applied; see Activating a manifest.The built-in default, which points at the reference examples:
config/examples/models.openrouter.yamlandconfig/examples/routing.minimal.yaml. Thealertsdefault isconfig/alerts.yaml, a path this repository does not ship — a missing alerts file means “use the built-in thresholds”.
Two things follow:
An environment variable beats the manifest. If you set
MODELS_CONFIG_PATHin a deployment that also has a manifest, the manifest’smodels:path is ignored, and the gateway logsexplicit env override ... wins over manifest value .... Pick one mechanism.Paths that come from the environment are resolved relative to the working directory of the process. Relative paths inside a manifest are resolved against the manifest file’s own directory.
A path that resolves from layer 1 or 2 but does not exist is a warning, not a
failure: the gateway logs Models config not found: <path> (source=env) and
continues with no models registered.
What a fresh clone does with no configuration at all
Start the gateway in a clean checkout with no MODELS_CONFIG_PATH, no manifest
and no overlay, and the built-in default applies: it serves
config/examples/models.openrouter.yaml. That file registers three models — two
straight OpenRouter routes and one that puts a local OpenAI-compatible server
first with OpenRouter as automatic fallback — and every one of them needs the
same single credential.
export OPENROUTER_API_KEY=sk-or-...
uv run uvicorn serving.servers.app:app --no-proxy-headers --port 8080
Without that variable every model in the file is skipped — its route’s only API key expands to empty — and the gateway says so before serving anything:
No models are available: every model in config/examples/models.openrouter.yaml was
skipped because its credential is unset. Set OPENROUTER_API_KEY and restart.
/v1/models will stay empty until then, and requests will report the model as not found.
This is the fastest way to a gateway that routes real traffic. Replace the file with your own registry once you know what you want to serve.
The model registry
Adding a New Model is the field reference. Three things belong here because they are properties of configuration loading rather than of any one field:
A route’s kind: picks the adapter; the model’s provider: is its default.
When a route omits kind:, the model’s top-level provider: is used; when a
model omits route: entirely, a single route is built from the top-level
provider:, base_url and api_key. provider is also the label written to
api_logs.provider and shown in metrics, which is why a route may override it
independently with provider:.
These are the built-in kinds. A deployment can add more with a backend extension; those are checked first.
Kind |
Adapter |
|---|---|
|
|
|
|
|
|
|
|
|
|
Local inference servers have no dedicated adapter: vllm, sglang and ollama
are OpenAI-compatible kinds that differ only in their provider label and usage
handling. Any other kind fails registry loading with
ValueError: Unknown adapter kind: <kind>.
Environment interpolation in the registry is whole-value only. In
models.yaml, only base_url, api_key, api_keys, provider_model_id and a
route’s embeddings_path are expanded, and only when the entire value is
exactly ${VAR}. There is no ${VAR:-default} and no embedded substitution:
base_url: ${LOCAL_BASE_URL} # expanded
base_url: ${LOCAL_BASE_URL:-http://x/v1} # NOT expanded — treated as a var
# named "LOCAL_BASE_URL:-http://x/v1",
# resolves empty, model is skipped
base_url: http://${LOCAL_HOST}/v1 # NOT expanded — the literal string,
# including "${LOCAL_HOST}", is used
The routing and alerts files are more permissive (see below), so do not carry habits from one file to the other.
The routing file
routing.yaml is optional, and most deployments can leave it out. Read this
section before relying on it: on a gateway with a database, which includes the
standard Docker stack, its weight split has no effect, and its health probe
never changes routing. Where traffic goes is decided by the weights in the
model registry, any weight overrides set in the admin console, and the router
described in Routing.
Every key, with its default:
Key |
Default |
Meaning |
|---|---|---|
|
|
The router for models that do not name one with |
|
|
Seconds a health probe waits for a connection. |
|
|
Seconds between health probes; |
|
|
Endpoints treated as local. |
|
|
Endpoints treated as remote. |
|
|
Free-form mapping. |
Each deployment entry needs both an endpoint: (which must start with http://
or https://) and a non-empty models: list; an entry with an empty models:
fails validation for the whole file, and the gateway logs
RoutingManager failed to initialize and carries on with the registry’s own
weights. An entry whose endpoint: expands to empty is dropped individually,
with a warning naming the orphaned models, so one unset variable cannot
invalidate every other endpoint.
default_router: fixed
timeout: 2
health_check: 30
local_deployment:
- endpoint: ${LOCAL_BASE_URL:-http://localhost:8000}
models: [<model-id>]
remote_deployment:
- endpoint: https://api.your-provider.example/v1
models: [<model-id>]
The weight split applies only to a gateway without a database. At startup,
the gateway sorts each model’s routes into local and remote: a route joins a
group when its base_url matches that entry’s endpoint: exactly and the
model is named in that entry’s models:. It then gives each group half the
traffic (unless the deprecated routing_parameter.local_fraction sets another
share), spreads that evenly within the group, and scales a model’s weights to
sum to 1.0. When one group is empty for a model, the other gets everything. A
model that matches nothing keeps the weights from its route: list. On a
gateway with a database, routing reads the registry’s weights plus any admin
overrides and never sees this split.
Health probing covers local_deployment endpoints only — remote providers do not
serve the gateway’s /health path and would be marked unhealthy for it. Each
probe is a GET to the endpoint’s origin root plus /health.
The probe only reports; it never changes routing. An endpoint that fails
every probe keeps its weight and keeps receiving traffic. Each change of state
is logged at WARNING, and GET /routing shows the latest results under
manager_status.endpoint_health, next to endpoint_health_enforced: false.
The circuit breaker is what
takes a failing endpoint out of rotation.
Unlike the model registry, this file’s expander handles ${VAR},
${VAR:-default}, and variables embedded in longer strings, at any depth.
routing_strategy: and routing_parameter: are the deprecated spellings of
default_router: and per-model router_params:. They still load, and log a
deprecation warning. Per-model router selection (router: / router_params: in
the model registry) overrides default_router and is documented in
Routing.
Settings that currently have no effect
The gateway accepts these when it loads its configuration, so setting them does not stop it from starting, but today they do nothing — or, in the case of the routing file’s weight split, nothing on a gateway with a database.
Setting |
Where |
What happens |
|---|---|---|
|
routing file |
Applied only on a gateway without a database; see The routing file |
|
routing file |
Logged and reported on |
|
model registry, |
Accepted and validated, not read |
|
distribution manifest |
Published in |
|
distribution manifest |
Accepted, not read by any component |
|
distribution manifest |
Accepted, not loaded into the terms page; use a Site UI module |
|
distribution manifest |
A label only; does not select a host or build anything |
Runtime configuration from the admin console
The files above are read once, at startup. Everything else an operator changes
about routing goes through the admin console — or the /admin/* endpoints
behind it — and is written to the operational store, so it takes effect without
a restart and survives one. This is the layer to use when a model, a provider or
a key has to exist now, on a running gateway, without editing the overlay and
redeploying.
It needs a database. DB_ENABLED defaults to true (Database
covers the connection); with DB_ENABLED=false there is no operational store,
every endpoint below answers 500 Database not configured, and the console
tabs that call them have nothing to write to. Every endpoint requires an
administrator’s JWT or ADMIN_TOKEN in the Authorization: Bearer ... header.
Providers tab
What |
Console |
Endpoint |
Stored in |
|---|---|---|---|
Add a custom OpenAI-compatible provider, with its first key |
Overview → Add provider |
|
|
Edit or remove a custom provider |
Overview → Edit provider |
|
|
Add a provider API key |
Keys → Add a new key |
|
|
Disable, re-enable or delete a key |
Keys |
|
|
Reserve a key for a tier |
Keys → Reserved for |
|
|
Take a provider out of rotation |
Availability → Enabled |
|
|
The registry on the Overview tab lists every provider the gateway can route to,
but only custom providers are editable there. A provider that code or the
model registry already owns — the built-in adapter kinds, and any provider:
label a route declares — is read-only in this table, and a stored definition
that reuses one of those slugs is skipped at boot rather than allowed to shadow
it. Custom providers are openai_compat only; a protocol the generic adapter
cannot speak needs an adapter, which is a code change
(Writing a Provider Adapter).
A key added here joins the same pool as the keys the registry names through
${VAR} and rotates with them. An environment key has no row of its own, so
the console can disable it or reserve it for a tier but cannot delete it; unset
the variable and restart for that.
Routing tab
What |
Console |
Endpoint |
Stored in |
|---|---|---|---|
Create a model that is not in the registry |
Create model |
|
|
Add a route to an existing model |
Add provider route |
|
|
Point a registry route somewhere else |
edit the route’s Target |
|
|
Change a route’s weight |
the weight field on each route |
|
|
Switch a model between |
the strategy selector |
|
|
Reserve a route for requests left waiting in the outbound queue or for an engine’s first token |
Queue offload |
|
|
Tune RouteWise for one model |
RouteWise settings |
|
|
GET /admin/routing/provider-routes (or .../{model_id}) returns every route
the gateway is serving, each tagged source: yaml, override or runtime,
and is the quickest way to see what the two layers add up to.
The provider selector on this tab offers every custom provider, plus each
built-in kind that the registry, the live route table or a configured
credential already names. Which route types a provider may be added as — on_demand,
quota or concurrency — is the deployment’s contract with that vendor and
is declared with PROVIDER_ROUTE_TYPES (see
Environment variables); an unlisted provider may
use any of the three.
A model created here is a runtime model: it exists only in the database,
carries an id, a first route, a pricing table, a router strategy and a
required_role (default admin, so a new model stays invisible to ordinary
users until you lower it), and is restored at every start. It does not carry
the catalog metadata a registry entry declares — context length, modalities,
supported parameters, aliases — and takes ModelConfig’s defaults for those:
context_length 8192, max_output_length 4096, text in and out, no tool
support. When a model needs any of that, put it in the registry.
Settings tab
What |
Console |
Endpoint |
Stored in |
|---|---|---|---|
Change who can see a model |
Model Visibility |
|
|
Exempt a model from the per-user concurrency limit |
Model Concurrency Limit |
|
|
How the two layers combine
The registry is loaded first and the stored state is applied on top of it, in this order at boot; each admin change is also applied to the running router the moment it is saved.
Custom provider definitions, skipping any slug the code or registry owns.
Stored provider keys, seeded into each provider’s pool beside the environment keys.
Per-model router strategy overrides.
Runtime routes. A runtime model is restored only if its creation finished; a stored route whose model is no longer in the registry is skipped with a warning, so removing a model from YAML does not bring it back through a leftover row.
Route overrides, retargeting the registry routes they name.
Key tier reservations, re-read now that every provider is known; then weight overrides, disabled providers, offload routes and model visibility, loaded into resolvers the router consults at request time.
Two rules follow from that order. A stored change never edits the file it
overrides: delete the override and the registry route, weight or router:
value is back, and a runtime candidate sits beside the registry’s routes rather
than replacing them. And the file still wins for anything the store has no row
for, so a registry edit plus restart is how catalog metadata, aliases and new
adapter kinds change.
Environment variables
The gateway reads a .env file from its working directory and the process
environment, case-insensitively; the process environment wins. .env.example
in the repository root is the annotated list — copy it to .env and edit.
Installation lists the variables you
are most likely to set. Secrets belong here and only here: not in the model
registry (reference them as ${VAR}), not in the manifest.
Running with your configuration
uv run uvicorn serving.servers.app:app --no-proxy-headers --port 8080
Then check what actually loaded:
curl -s localhost:8080/health # includes routes_configured
curl -s localhost:8080/v1/models # generated from the registered adapters
The startup log says what was loaded. Registered N routes from <path> names
the registry file that was used; [distribution dark mode] ... lines show what
a manifest would change once activated; Skipping model '<id>' ... after env expansion names each model whose credentials were unset.
Warning
GET /routing requires no credentials and returns, for every published route,
the upstream base_url, its provider label and its weight — that is, your
complete upstream topology including any host and port embedded in a base URL.
Block /routing at your reverse proxy on any gateway reachable from the
internet unless you intend to publish this topology. GET /admin/routing
additionally reports unpublished routes and requires an administrator’s JWT
or ADMIN_TOKEN in the Authorization: Bearer ... header. Treat both endpoints’
output as sensitive.