Backend Extensions
A backend extension is a trusted Python module that a deployment loads into the gateway at startup. It can add an adapter kind, decide who may use the Cloud Agent, or report provider quotas. This page is for developers writing one; for configuration-only changes, see Distribution customization.
Loading an extension
BACKEND_EXTENSIONS is an optional comma-delimited list of trusted, local Python
module names. It is empty by default. Each module must expose a synchronous
register() function that takes no arguments. The gateway imports and registers
each module once per process, after loading dotenv and before constructing
runtime routes or consuming their registries. An import or registration failure
aborts startup; the gateway does not silently use a different adapter.
Extensions register factories through
serving.servers.registry.register_adapter_factory(kind, factory, *, override=False).
A factory receives a configuration dictionary and returns an adapter, before
built-in provider defaults are applied. Registering a kind that another
extension has already registered is an error. Replacing a built-in kind requires override=True and emits a log
entry. Registered kinds also become reserved provider labels. The startup
register() function may populate the existing runtime-setting and provider
metadata dictionaries before their consumers run; it must mutate those shared
dictionaries rather than replace them.
The deployment must make these modules importable, for example through its read-only overlay mount. This executes trusted server code, not user-supplied configuration: do not derive module names from requests or allow an admin form to choose them. The loader does not fetch remote modules or load code per request. Import all extension code during startup, avoid later lazy imports or live code reloads, and restart the backend when changing the mounted extension. See Deployment-local adapters for a minimal factory example.
Cloud Agent access
The gateway keeps its fixed role hierarchy and resolves current account state.
The deployment owns the rule for who can use or administer its Cloud Agent.
A trusted extension registers one synchronous callback during register() with
serving.agent_access.register_agent_access_policy(policy). The callback receives
a read-only mapping containing exactly user_id and role, after the gateway
has verified that the user exists and has status active. It receives no
profile, credentials, or database connection.
The callback returns a list[str] or tuple[str, ...] of explicit permissions:
agent.use permits Agent use, and agent.admin permits Agent administration.
Administration requires both permissions. Unknown, repeated, or non-string
permissions are invalid; values are not coerced. Registering a second policy is
an error. With no registered policy, Agent access is denied for every role.
The gateway imposes no deployment-specific role-to-permission rule.
The standalone Agent reads GET /internal/users/{user_id}/agent-access using
the existing GATEWAY_GRANT_DISPATCH_TOKEN bearer token. A successful response
contains exactly user_id, allowed, and permissions; allowed is true
exactly when agent.use is present. Permissions are returned in the order
agent.use, agent.admin. An intentional denial is HTTP 200 with
allowed: false and permissions: []. The endpoint reads the gateway’s
operational store and evaluates the policy on each request; it does not cache
permission decisions. Account details can be up to 60 seconds old: a change
made through this gateway process is seen at once, but one made directly in the
database or by another process is seen when the cached account expires.
Unknown or inactive accounts receive HTTP 403 with error type
subject_unavailable, before the callback runs. A callback exception or invalid
result receives HTTP 503 with error type agent_access_unavailable, without
policy details. Both use the gateway’s standard {"error": {"type": ..., "message": ...}} envelope. A request without a valid dispatch token gets 401,
and a gateway with the internal API turned off answers 404.
Quota reporting
The admin console’s Providers → Quotas tab and RouteWise’s quota-aware routing both read usage from quota sources. The gateway ships with none: a backend extension registers one per provider, reading the provider’s usage API or your own metering service. Nothing requires a particular supplier, website login or cookie.
An extension’s register() calls
serving.admin.provider_quotas.register_quota_fetcher(provider, display_name, fetch);
fetch is awaited as fetch(operational_store, services) and returns one
ProviderQuotaResult per configured key, converting its own failures into
results rather than raising. The admin endpoint supplies the store and services;
RouteWise currently calls it with (None, None), so the fetcher must also work
without those objects. Registration runs at startup, not when the module is
merely imported. A fetcher queries usage; it does not send user inference
requests or change the route’s inference protocol.
Each result contains ProviderQuotaUsage rows with label, used, limit,
unit and an optional timezone-aware reset_at. The deployment’s provider
identifier links the source to its routes; it is not a hard-coded vendor
enum. quota_source.provider, usage_label and unit must exactly match a
returned result and usage row. Report the account/window shared by that quota
pool, and keep independent accounts in distinct sources. Do not sum unrelated
keys or present a failed query as zero usage.
The UI can display percentages, currencies and other units, but RouteWise’s
quota admission consumes one request at a time and requires a compatible
count-based source. A percentage alone is not a request allowance. The source
owns window boundaries and reset times; quota.limit is a cross-check against
its reported limit, not a replacement measurement. A source with no successful
snapshot stays unready. A failed refresh does not manufacture a fresh balance;
an existing pool retains its last successful snapshot and local increments.
Run a local quota source
The complete example extension
reads /usage from the bundled fake provider. Its in-memory daily counter is
for demonstration only: it resets on process restart or at UTC midnight, and is
not production accounting. It uses no real account or credential and is loaded
only when explicitly selected through BACKEND_EXTENSIONS.
After the developer setup, run these commands from the repository root in three terminals. First, start a simulated quota provider:
uv run python distributions/example/fixtures/fake-openai-provider/server.py \
--port 18353 --response-text ROUTED_TO_QUOTA --quota-limit 100
Then start a slower, priced fallback (its prices in the example are fictional):
uv run python distributions/example/fixtures/fake-openai-provider/server.py \
--port 18352 --response-text ROUTED_TO_FALLBACK --ttft-delay-ms 400
Finally start the gateway with the opt-in quota registry. Name the demo token once so the admin call below can reuse it:
export DEMO_ADMIN_TOKEN=local-quota-demo-only
PYTHONPATH=.:apps/backend \
PYTHON_DOTENV_DISABLED=1 \
BACKEND_EXTENSIONS=distributions.example.quota_extension \
EXAMPLE_QUOTA_BASE_URL=http://127.0.0.1:18353 \
MODELS_CONFIG_PATH=config/examples/models.routewise.quota.yaml \
ROUTING_CONFIG_PATH=config/examples/routing.minimal.yaml \
DB_ENABLED=false USER_AUTH_ENABLED=false ADMIN_TOKEN="${DEMO_ADMIN_TOKEN}" \
JWT_SECRET_KEY=local-quota-demo-signing-secret-not-for-production \
uv run uvicorn serving.servers.app:app --no-proxy-headers --host 127.0.0.1 --port 18080
These settings disable dotenv loading and user authentication and use a public demo-only admin token and JWT signing secret: keep the listener on loopback and never use them for a real deployment. The admin endpoint needs the signing secret even when authenticating with the demo token. Read the same result the console uses:
curl http://127.0.0.1:18080/admin/provider-quotas \
-H "Authorization: Bearer ${DEMO_ADMIN_TOKEN}"
Send several requests to calibrate the cost envelope and allow the first five-second probe/snapshot cycle to finish:
curl http://127.0.0.1:18080/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"quota-demo","messages":[{"role":"user","content":"hi"}]}'
The response identifies the selected fixture. Initial requests use the fallback
until the quota route has both a usage snapshot and cost-envelope evidence.
Accepted chat requests, including active probes, consume the mock quota;
health, model discovery and /usage reads do not. To simulate exhaustion,
restart only the quota fixture with --quota-limit 100 --quota-used 100; after
the next snapshot refresh, requests continue through the fallback. Stop each
local process with Ctrl+C when finished.
To see the card in the authenticated Web/Admin Console, configure this same
extension and registry on a local full-stack demo backend, then open
Providers → Quotas. Use a quota endpoint reachable from that backend:
inside Compose, 127.0.0.1 denotes the backend container, not the host. The
Quickstart covers the full-stack demo and its
authentication. With no returned results the Quotas tab stays visible and
links back here; the Overview, Keys, Availability and Performance tabs work
without any quota source.
For your own integration, copy the extension and replace its local HTTP query and mapping with your authorized data source. Preserve the result contract, bound network requests with a timeout, handle failures explicitly and keep credentials out of responses and logs. The provider’s own terms still govern how you may query its usage.