Deployment Guide
Running HybridInference as a long-lived deployment: what starts, how to reach it publicly or keep it private, how to create the first administrator, how to operate it, and how to reset it without losing (or accidentally keeping) data. A staging or scratch instance is the same stack; the sections on keeping it private and on the first admin account matter most there.
First-time setup — cloning, filling in .env, and the first make up — is in
Installation. This page assumes the
stack already comes up.
What the stack is
make up starts three containers from deploy/docker/docker-compose.yml:
Service |
Image / build |
Published on |
|---|---|---|
|
built from |
|
|
built from |
|
|
|
|
All three published ports default to loopback. A reverse proxy on the same
host can reach the console at 127.0.0.1:3001. The containers also join one bridge
network defined in the same Compose file, on which the backend reaches the
database as postgres:5432. Of the database address settings, only DB_PORT
from .env reaches this file, and only as the host half of the mapping (127.0.0.1:${DB_PORT:-5432}:5432); the
host-side bind address is hard-coded to loopback. DB_HOST is pinned to
postgres in the Compose file and is ignored under Compose — it matters only
for a backend started directly from source.
One more service exists in the same file but starts only when its profile is
named: pgadmin (profile admin).
frontend depends on backend with condition: service_started, not service_healthy — deliberately, so that a backend
reporting unhealthy because its database logging is down does not stop the
console from starting.
Putting it on the public internet
Nothing in the stack terminates TLS, and this repository ships no reverse-proxy
config to copy: certificates and the proxy in front of the two published ports
are yours to supply. A proxy running on the host can use 127.0.0.1:3001 for
the console or 127.0.0.1:8080 for direct gateway access. A proxy container on
the same Docker network can use frontend:3001 or backend:8080.
If your proxy runs on another machine, set FRONTEND_HOST to a reachable host
interface in .env; 0.0.0.0 binds all IPv4 interfaces. Restrict access to the
intended proxy and run make up to recreate the port mapping. The console
forwards API routes as well as serving pages, so exposing its port also exposes
those routes. Set BACKEND_HOST separately only if direct gateway access is
needed.
Releases before the loopback default bound the console to 0.0.0.0. If a
deployment relied on that, set FRONTEND_HOST explicitly before upgrading;
see Releases and upgrades.
Which public paths the console serves itself and which it forwards to the
backend is a separate question, and the answer is in the console’s own
next.config.js rather than in any proxy config. See
The public path table.
What your proxy still has to do
Three protections are commonly left to the reverse proxy, and nothing inside
this stack provides them. A deployment that exposes the published ports without
a proxy — a tunnel daemon connecting to 127.0.0.1:3001, for instance — has
none of them, and nothing warns you:
Protection |
What a proxy in front typically does |
|---|---|
Next.js Server Action guard |
Refuse requests carrying a |
Request body cap on |
Bound completion request bodies (commonly |
|
Overwrite a client-supplied chain, so only hops the proxy inserted reach the gateway |
The gateway reads forwarded headers only from the proxies you authorize (see
Trusted proxies and client IPs), but it
never rewrites them. Without a proxy that does, the leftmost X-Forwarded-For
entry is whatever the caller sent — which is why CF-Connecting-IP is
preferred when Cloudflare is in front.
Keeping it private: an SSH tunnel
A staging or scratch instance, or any gateway you do not want on a network at all, can keep the loopback defaults. Reach it from your workstation by forwarding both ports:
ssh -L 3001:127.0.0.1:3001 -L 8080:127.0.0.1:8080 <user>@<your-server>
Then open http://localhost:3001. Keep the local end on localhost or
127.0.0.1: the refresh cookie is issued with the Secure flag by default
(COOKIE_SECURE), and browsers accept Secure cookies only over HTTPS or from
a loopback origin.
The console talks to the API at whatever NEXT_PUBLIC_API_BASE was baked in at
image build time — http://localhost:8080 by default — which is why the tunnel
forwards 8080 as well. Changing it is a frontend rebuild, not a restart.
VS Code-family editors can manage the same forwards from their ports panel.
Check the stack answers:
curl -s http://localhost:8080/health
curl -s http://localhost:8080/v1/models
The default CORS_ALLOWED_ORIGINS already covers ports 3000, 3001 and 3002 on
localhost and 127.0.0.1, plus HTTPS on port 8443, so a tunnelled instance
needs no CORS entry. Add one only when you serve the console from
another origin.
Everyday operations
All from the repository root:
make up # start everything
make down # stop everything (data survives; see below)
make restart # restart everything
make restart s=backend # restart one service
make ps # services and health status
make logs # tail all logs
make logs s=backend # tail one service
make build # rebuild images and restart
make build s=frontend # rebuild one service
make up and make build first run the docker-volumes target, which creates
the external volume hybridinference_postgres_data when it is missing.
To start an optional profile, pass it on the make command line — a variable
set there is exported into the environment of the recipe, and a shell variable
outranks every --env-file in Compose:
make up COMPOSE_PROFILES=admin
COMPOSE_PROFILES is a comma-separated list, so you can name several profiles
at once. It can also be set in .env (as .env.example notes), but the
command line is the form to reach for when you want certainty about which
profiles are active.
What a change actually requires
Three different answers, and picking the wrong one looks like the change not taking effect:
You changed |
Do this |
|---|---|
A value in |
|
A model registry or routing YAML |
|
A distribution branding YAML |
|
A file in the mounted site-assets directory |
No image rebuild; replace the file in the deployment overlay |
|
|
A true build-only |
|
Backend or frontend source |
|
The console gets its name and branding from the backend’s /site-config at
runtime, and the /agents destinations from its own runtime environment, so
neither needs a new frontend image: restart the backend after editing the
branding file, and run make up after changing the console’s environment.
The console will not render its normal pages without a valid /site-config
answer. If the backend is unreachable, slow (over three seconds) or returns
something invalid, the console shows a configuration error with a retry button
instead of guessing. Check the frontend logs and whether the backend is up,
then retry. The example deployment needs no extra settings for this.
A few older build-time options remain for existing build pipelines: the old
branding variables, which Compose still accepts as build arguments, and two
agent build arguments that it no longer sets. Standard builds do not need them.
Anything that really is a NEXT_PUBLIC_* build value is compiled into the
browser bundle and needs make build s=frontend; see
The public path table.
Configuration
Environment
Everything is in .env at the repository root; .env.example is the annotated
list. Compose is invoked with --env-file .env and the backend service also
loads it as env_file. The variables Compose itself requires, and the two
secrets you should not leave blank, are in
Installation.
Config file resolution
The gateway picks each configuration file from an explicit
MODELS_CONFIG_PATH / ROUTING_CONFIG_PATH / ALERTS_CONFIG_PATH, then an
active distribution manifest, then the examples under config/examples/; see
How a gateway finds its config.
Note that the Compose file passes these through explicitly:
ROUTING_CONFIG_PATH: ${ROUTING_CONFIG_PATH-}
MODELS_CONFIG_PATH: ${MODELS_CONFIG_PATH-}
DISTRIBUTION_CONFIG_PATH: ${DISTRIBUTION_CONFIG_PATH-}
An --env-file alone does not put a variable into a container’s environment;
these lines are what carry it in. Without them the gateway would silently fall
back to the default files.
Local inference servers
The backend container reaches servers on the host through
host.docker.internal, which the Compose file wires with
extra_hosts: host.docker.internal:host-gateway. Write that address explicitly
in the model registry:
route:
- kind: openai_compat
base_url: http://host.docker.internal:8001/v1
For a backend running directly on the host, use localhost instead. The gateway
never rewrites provider URLs. See
Adding a New Local Model.
Health checks
curl -s http://localhost:8080/health
{
"status": "healthy",
"routes_configured": 3,
"database_configured": true,
"database_connected": true,
"stores": {
"operational_store": {"status": "ok", "backend": "postgres", "cache": "in_memory"},
"log_store": {"status": "ok", "backend": "postgres"}
}
}
routes_configuredcounts published route entries — one per model id in the active registry, plus one per alias. It is whatever your registry defines.database_configureddistinguishes “this deployment asked for no database” from “the database is down”: withDB_ENABLED=falseit isfalseand the status is stillhealthy; with a database configured but unreachable at startup,/healthanswers 503 with"reason": "database_unavailable_at_startup".statusbecomesdegraded— still HTTP 200 — when one configured store is down but the other is serving. That shape is deliberate: the containerHEALTHCHECKusescurl -f /health, so returning 503 for partial degradation would tear down backends that are still answering requests.
Use /health/ready for a strict readiness probe: it applies AND-logic across
configured stores and returns 503 unless every one of them is up.
/health/deep additionally reports per-endpoint health.
Marking monitor traffic
A monitor that drives real inference — hitting /v1/chat/completions on a
schedule to measure a backend end to end, rather than just polling /health —
would otherwise land in api_logs, skew the dashboards, and feed RouteWise’s
online learning as if it were user demand. Send X-Probe: synthetic on those
requests to keep them out: a marked request is left out of api_logs (and its
rejections out of the rejection log), does not record a routing observation,
and carries an X-Provider response header naming the backend that answered,
so the monitor can confirm which route it exercised.
The marker is honoured only from an authenticated internal- or admin-role API
key — never a free/pro key, an agent-sandbox grant, or, importantly, an
anonymous caller on a deployment running with USER_AUTH_ENABLED=0 (auth-off
hands every caller the admin role, which is not the same as holding a monitor
identity). From any other caller the header is ignored and the request is
logged like ordinary traffic. A deployment that needs probes without auth wants
an explicit mechanism — a shared secret, a source allowlist — not this header.
The marker is about noise, not access. It cannot keep a monitor’s own auth failures from tripping the repeated-auth-failure blocklist, because that decision is made before the presented key is read — see A monitor or service account is suddenly getting 429s.
The one thing the marker never touches is billing: cost and quota are
incremented unconditionally on every surface, for trusted and untrusted callers
alike, so a probe cannot be used to obtain unmetered inference. To keep marked
traffic in api_logs after all — to see a monitor’s real latency and spend in
the requests dashboard — turn on the log_synthetic_probes runtime setting;
the routing-observation and X-Provider behaviour is unchanged.
Alerting
The backend has an in-process alert engine that posts to a Slack webhook. It is off unless you turn it on:
ALERTS_ENABLED=true
SLACK_ALERTS_WEBHOOK_URL=https://hooks.slack.com/services/...
If SLACK_ALERTS_WEBHOOK_URL is empty it falls back to SLACK_WEBHOOK_URL, so
one webhook can serve both code paths.
Rules and thresholds are a deployment’s own; this repository ships no alerts
file. Point ALERTS_CONFIG_PATH at yours, or leave it unset and the built-in
thresholds apply. The rule types and evaluation live in
apps/backend/serving/observability/.
Two auth rules are worth knowing apart, because their defaults differ on purpose:
rules.auth_failure_spikeis off. Bad keys are internet background noise, and a count of them names nothing to act on. Theauth_failurelog records are emitted regardless.rules.auth_ip_blockedis on. This one fires when the blocklist starts refusing a source — a discrete decision at a much higher threshold, naming an address. It is on because the source is sometimes the deployment’s own; see the troubleshooting entry below.
What an auth alert tells you
An auth-failure card names the sources and, where it can, the accounts behind them — the addresses they came from, how many distinct ones, the leading characters of the keys presented, why each failed, and the paths being hit. The recovery card carries the same picture of the incident that just closed, rather than only the rule’s name: by the time a spike resolves, the window it breached on is empty, so the numbers have to come from a tally kept across the incident.
Two of those lines are worth reading carefully:
Known accounts. Most auth failures are anonymous by construction — nobody was authenticated, which is the failure. A named account means a key this deployment did issue was presented and refused, with
credential_statesaying why (revoked,expired,user_suspended). That is the actionable case: a monitor, CI job or service account whose credential went stale. Resolving the owner costs one indexed lookup per failed auth, bounded by the shared rejection-enrichment budget and shed instantly under a flood; setAUTH_FAILURE_IDENTIFY_CALLER=falseto spend nothing and lose the line. Read a named account as evidence and no named account as unknown, never as proof the traffic is external: the lookup is shed during exactly the flood you are investigating, and answers nothing on a timeout, a failed lookup, or with the setting off.Arrived via peers. Present only when the reported addresses did not come off the socket. They are then only as trustworthy as the proxy that set them, and a forged
X-Forwarded-Foris exactly how a source spreads its failures across the blocklist’s buckets. The line names the sockets they actually arrived on.
Counts marked (capped) are floors, not totals: a source rotating addresses
faster than the tally tracks them stops being counted rather than being allowed
to grow it without bound. Do not size an incident from a capped number.
Silencing alerts from the dashboard
The admin dashboard’s Settings tab has two controls, and neither changes your alerts file:
Slack Alerts snoozes every alert for up to seven days.
Alert Types mutes one type of alert, such as
auth_ip_blockedorcircuit_open, for an hour, a day, a week, or until it is unmuted. Every other type keeps sending.
A mute is stored in the database, so it survives a restart and reaches every gateway process within a few seconds. While a type is muted nothing of it is sent, and an incident that opens during the mute also closes without a recovery message. An incident that was posted before the mute still gets its recovery, and a breach still live when the mute lifts pages at its next evaluation. A mute is for an alert you want back later; to retire a rule for good, turn it off in your alerts file instead.
The first admin account
Every deployment needs one administrator to start with. Choose how you create it with care:
Warning
Do not combine SIGNUP_ENABLED=1, SIGNUP_REQUIRE_EMAIL_VERIFICATION=0 and
ADMIN_EMAILS on an instance anyone else can reach. Together they are a
privilege-escalation recipe:
with verification disabled,
POST /auth/signupmarks any address as verified without sending mail to it;on login and on every token refresh, the backend promotes any account whose address is listed in
ADMIN_EMAILSfromfreetoadmin, with no check that the person signing up owns that address.
So a stranger who guesses or reads your ADMIN_EMAILS value signs up with that
address and is an admin on their first login.
Pick one of these instead.
Preferred — create the admin out of band and leave ADMIN_EMAILS unset.
ops/admin/create_admin.py writes the row directly: it creates the account with
role='admin', status='active', email_verified=TRUE, or promotes an
existing account with the same address. Run it from the repository root once the
backend has started at least once (the backend creates the schema):
python ops/admin/create_admin.py --email [email protected]
Run it inside the project environment (source .venv/bin/activate after
make setup-dev) so the serving package is importable. It reads DB_HOST,
DB_PORT, DB_NAME, DB_USER and DB_PASSWORD from .env, and Postgres
publishes on 127.0.0.1:5432, so it works from the host shell. Omit
--password and it prompts, keeping the password out of your shell history.
With this in place the instance can run with signup closed:
USER_AUTH_ENABLED=1
SIGNUP_ENABLED=0
Alternative — keep signup open, but leave verification on.
SIGNUP_REQUIRE_EMAIL_VERIFICATION defaults to true, and with it on an
account cannot log in until it has followed a link sent to the address, which
restores the ownership check that ADMIN_EMAILS itself does not perform. This
needs working SMTP; without it nobody can complete a signup.
ADMIN_EMAILS also picks the default recipients for signup approval mail. To
narrow the notification list without changing who holds the admin role, set
SIGNUP_NOTIFY_EMAILS (comma-separated); when it is empty, notifications fall
back to ADMIN_EMAILS.
SIGNUP_NOTIFY_EMAILS=[email protected]
Both signup_enabled and signup_require_email_verification can also be
flipped at runtime through the settings store, and the runtime value wins over
the environment. A .env line is the starting point, not a guarantee.
Database
PostgreSQL 16 runs in the postgres service with its data in the Docker volume
hybridinference_postgres_data. A psql shell:
docker exec -it hybridinference-postgres psql -U "${DB_USER}" -d "${DB_NAME}"
Schema details are in Database.
pgAdmin (optional)
make up COMPOSE_PROFILES=admin
pgAdmin then listens on 127.0.0.1:5050 with SCRIPT_NAME=/pgadmin, so an SSH
tunnel to that port is enough to reach it. The console can also proxy it at
/pgadmin/, gated on an admin session by
apps/frontend/src/app/pgadmin/[[...path]]/route.ts, which denies on every
unexpected condition, including a backend it cannot reach.
One thing to get right: whether pgAdmin also asks for its own login is set by
PGADMIN_CONFIG_SERVER_MODE, and the two defaults disagree. The Compose service
falls back to False, which serves pgAdmin with no login at all; .env.example
suggests True, which turns pgAdmin’s own login on behind the console’s gate.
True is the safer of the two.
Troubleshooting
A service will not start
make logs s=backend
make ps
required variable DB_NAME is missing a value: DB_NAME must be set in .env file— Compose stopped at interpolation before starting anything.DB_NAME,DB_USERandDB_PASSWORDare declared required with the${VAR:?message}form, so the half after the colon is the Compose file’s own text and the most greppable part of the line.Port already in use — override
BACKEND_PORT,FRONTEND_PORTorDB_PORT.Database connection failed — check
make psfor thepostgreshealth status.
A monitor or service account is suddenly getting 429s
Too many authentication failures from this IP. Temporarily blocked. is the
gateway’s own abuse defense, not a provider error and not a quota. Once a source
accumulates AUTH_FAILURE_BLOCK_THRESHOLD failed authentications inside
AUTH_FAILURE_BLOCK_WINDOW_SEC (200 in a day, by default) it is refused for
AUTH_FAILURE_BLOCK_DURATION_SEC (a day).
The awkward case is a caller you own — a status monitor, a CI job, a service account — whose key was rotated, revoked, or never reached its environment. It retries on a schedule, crosses the threshold, and is then refused ahead of the key check, which has two consequences:
Repairing the credential does not lift the block. The blocklist is consulted before the presented key is read, so a corrected key gets the same 429 until the deadline passes.
The 429 hides the original error. Whatever the caller reports after the block is in place says nothing about whether the underlying 401/403 was fixed.
Which caller is it? Turn on the log_rejected_requests admin setting and the
refusals land in Recent Requests as ip_blocked rows. Each row names the
account behind the key the caller presented — including a key that was revoked
or expired, which is what a stuck monitor is presenting — and labels it with the
credential’s state (revoked, expired, user_suspended) beside the user. A
row with no user is unresolved, not proof of a stranger: the key may be one this
deployment never issued, or the lookup may have been shed, since it runs on a
strict budget that gives up first under exactly the flood a block is holding
back.
To recover, first fix the credential, then clear the block:
# Which sources is this worker refusing?
curl -s -H "Authorization: Bearer $ADMIN_TOKEN" \
http://localhost:8080/admin/auth-blocks
# Lift one. `ip` takes a raw address, or a bucket key exactly as listed
# (IPv6 sources are bucketed to their /64).
curl -s -X POST -H "Authorization: Bearer $ADMIN_TOKEN" \
-H 'Content-Type: application/json' \
-d '{"ip": "203.0.113.7"}' \
http://localhost:8080/admin/auth-blocks/clear
cleared: false means there was nothing to lift — it lapsed, or that bucket was
never blocked. Clearing grants no immunity: a caller still presenting a bad key
is blocked again on crossing the threshold. For a source that should never be
blocked at all, list it in AUTH_FAILURE_BLOCK_EXEMPT_IPS (comma-separated
addresses or CIDRs) instead.
The blocklist is per-process, in-memory state. On the single-process default both endpoints are exact, and a restart also clears every block. Run multiple workers and each holds its own counts, so a listing shows only the worker that answered and clearing may take more than one call.
Resetting the stack
Stop and start, keeping data:
make down && make up
Destroy the database and start clean. docker compose down -v does not do
this. postgres_data is declared external: true in
deploy/docker/docker-compose.yml, and Compose never removes an external
volume — down -v returns success and leaves it fully intact, so make up
comes back on exactly the same data. Remove it by name:
make down
docker volume rm hybridinference_postgres_data
make up # docker-volumes recreates it empty; Postgres re-initialises
Warning
docker volume rm is irreversible and takes every account, API key and request
log with it. Take a pg_dump first if any of it matters.
pgAdmin’s own volume (hybridinference_pgadmin_data) is an ordinary local
volume that make down leaves in place. To clear pgAdmin’s saved state, remove
it by name as described in Resetting pgAdmin.
Rebuilding after code changes
make build # all images
make build s=backend # one service