Deployment Guide

Running HybridInference as a long-lived deployment: what starts, how to reach it publicly or keep it private, how to create the first administrator, how to operate it, and how to reset it without losing (or accidentally keeping) data. A staging or scratch instance is the same stack; the sections on keeping it private and on the first admin account matter most there.

First-time setup — cloning, filling in .env, and the first make up — is in Installation. This page assumes the stack already comes up.

What the stack is

make up starts three containers from deploy/docker/docker-compose.yml:

Service

Image / build

Published on

backend

built from deploy/docker/Dockerfile.backend

${BACKEND_HOST:-127.0.0.1}:${BACKEND_PORT:-8080}

frontend

built from deploy/docker/Dockerfile.frontend

${FRONTEND_HOST:-127.0.0.1}:${FRONTEND_PORT:-3001}

postgres

postgres:16

127.0.0.1:${DB_PORT:-5432}

All three published ports default to loopback. A reverse proxy on the same host can reach the console at 127.0.0.1:3001. The containers also join one bridge network defined in the same Compose file, on which the backend reaches the database as postgres:5432. Of the database address settings, only DB_PORT from .env reaches this file, and only as the host half of the mapping (127.0.0.1:${DB_PORT:-5432}:5432); the host-side bind address is hard-coded to loopback. DB_HOST is pinned to postgres in the Compose file and is ignored under Compose — it matters only for a backend started directly from source.

One more service exists in the same file but starts only when its profile is named: pgadmin (profile admin).

frontend depends on backend with condition: service_started, not service_healthy — deliberately, so that a backend reporting unhealthy because its database logging is down does not stop the console from starting.

Putting it on the public internet

Nothing in the stack terminates TLS, and this repository ships no reverse-proxy config to copy: certificates and the proxy in front of the two published ports are yours to supply. A proxy running on the host can use 127.0.0.1:3001 for the console or 127.0.0.1:8080 for direct gateway access. A proxy container on the same Docker network can use frontend:3001 or backend:8080.

If your proxy runs on another machine, set FRONTEND_HOST to a reachable host interface in .env; 0.0.0.0 binds all IPv4 interfaces. Restrict access to the intended proxy and run make up to recreate the port mapping. The console forwards API routes as well as serving pages, so exposing its port also exposes those routes. Set BACKEND_HOST separately only if direct gateway access is needed.

Releases before the loopback default bound the console to 0.0.0.0. If a deployment relied on that, set FRONTEND_HOST explicitly before upgrading; see Releases and upgrades.

Which public paths the console serves itself and which it forwards to the backend is a separate question, and the answer is in the console’s own next.config.js rather than in any proxy config. See The public path table.

What your proxy still has to do

Three protections are commonly left to the reverse proxy, and nothing inside this stack provides them. A deployment that exposes the published ports without a proxy — a tunnel daemon connecting to 127.0.0.1:3001, for instance — has none of them, and nothing warns you:

Protection

What a proxy in front typically does

Next.js Server Action guard

Refuse requests carrying a Next-Action header (if ($http_next_action) { return 403; }), so console server actions cannot be invoked from outside the console

Request body cap on /v1/

Bound completion request bodies (commonly client_max_body_size 50m). The gateway enforces no size limit of its own

X-Forwarded-For rewriting

Overwrite a client-supplied chain, so only hops the proxy inserted reach the gateway

The gateway reads forwarded headers only from the proxies you authorize (see Trusted proxies and client IPs), but it never rewrites them. Without a proxy that does, the leftmost X-Forwarded-For entry is whatever the caller sent — which is why CF-Connecting-IP is preferred when Cloudflare is in front.

Keeping it private: an SSH tunnel

A staging or scratch instance, or any gateway you do not want on a network at all, can keep the loopback defaults. Reach it from your workstation by forwarding both ports:

ssh -L 3001:127.0.0.1:3001 -L 8080:127.0.0.1:8080 <user>@<your-server>

Then open http://localhost:3001. Keep the local end on localhost or 127.0.0.1: the refresh cookie is issued with the Secure flag by default (COOKIE_SECURE), and browsers accept Secure cookies only over HTTPS or from a loopback origin.

The console talks to the API at whatever NEXT_PUBLIC_API_BASE was baked in at image build time — http://localhost:8080 by default — which is why the tunnel forwards 8080 as well. Changing it is a frontend rebuild, not a restart.

VS Code-family editors can manage the same forwards from their ports panel.

Check the stack answers:

curl -s http://localhost:8080/health
curl -s http://localhost:8080/v1/models

The default CORS_ALLOWED_ORIGINS already covers ports 3000, 3001 and 3002 on localhost and 127.0.0.1, plus HTTPS on port 8443, so a tunnelled instance needs no CORS entry. Add one only when you serve the console from another origin.

Everyday operations

All from the repository root:

make up                  # start everything
make down                # stop everything (data survives; see below)
make restart             # restart everything
make restart s=backend   # restart one service
make ps                  # services and health status
make logs                # tail all logs
make logs s=backend      # tail one service
make build               # rebuild images and restart
make build s=frontend    # rebuild one service

make up and make build first run the docker-volumes target, which creates the external volume hybridinference_postgres_data when it is missing.

To start an optional profile, pass it on the make command line — a variable set there is exported into the environment of the recipe, and a shell variable outranks every --env-file in Compose:

make up COMPOSE_PROFILES=admin

COMPOSE_PROFILES is a comma-separated list, so you can name several profiles at once. It can also be set in .env (as .env.example notes), but the command line is the form to reach for when you want certainty about which profiles are active.

What a change actually requires

Three different answers, and picking the wrong one looks like the change not taking effect:

You changed

Do this

A value in .env

make up — a container reads its env_file when it is created, so docker compose restart keeps the old environment

A model registry or routing YAML

make restart s=backend — config/ and distributions/ are bind-mounted read-only, so no rebuild is needed

A distribution branding YAML

make restart s=backend — /site-config serves the validated snapshot loaded at backend startup

A file in the mounted site-assets directory

No image rebuild; replace the file in the deployment overlay

AGENT_WEB_INTERNAL_URL or AGENT_CONTROL_PLANE_INTERNAL_URL

make up — the /agents route handler reads them at runtime in the recreated frontend container

A true build-only NEXT_PUBLIC_* compatibility value

make build s=frontend — see below

Backend or frontend source

make build, or make build s=<service>

The console gets its name and branding from the backend’s /site-config at runtime, and the /agents destinations from its own runtime environment, so neither needs a new frontend image: restart the backend after editing the branding file, and run make up after changing the console’s environment.

The console will not render its normal pages without a valid /site-config answer. If the backend is unreachable, slow (over three seconds) or returns something invalid, the console shows a configuration error with a retry button instead of guessing. Check the frontend logs and whether the backend is up, then retry. The example deployment needs no extra settings for this.

A few older build-time options remain for existing build pipelines: the old branding variables, which Compose still accepts as build arguments, and two agent build arguments that it no longer sets. Standard builds do not need them. Anything that really is a NEXT_PUBLIC_* build value is compiled into the browser bundle and needs make build s=frontend; see The public path table.

Configuration

Environment

Everything is in .env at the repository root; .env.example is the annotated list. Compose is invoked with --env-file .env and the backend service also loads it as env_file. The variables Compose itself requires, and the two secrets you should not leave blank, are in Installation.

Config file resolution

The gateway picks each configuration file from an explicit MODELS_CONFIG_PATH / ROUTING_CONFIG_PATH / ALERTS_CONFIG_PATH, then an active distribution manifest, then the examples under config/examples/; see How a gateway finds its config.

Note that the Compose file passes these through explicitly:

ROUTING_CONFIG_PATH: ${ROUTING_CONFIG_PATH-}
MODELS_CONFIG_PATH: ${MODELS_CONFIG_PATH-}
DISTRIBUTION_CONFIG_PATH: ${DISTRIBUTION_CONFIG_PATH-}

An --env-file alone does not put a variable into a container’s environment; these lines are what carry it in. Without them the gateway would silently fall back to the default files.

Local inference servers

The backend container reaches servers on the host through host.docker.internal, which the Compose file wires with extra_hosts: host.docker.internal:host-gateway. Write that address explicitly in the model registry:

route:
  - kind: openai_compat
    base_url: http://host.docker.internal:8001/v1

For a backend running directly on the host, use localhost instead. The gateway never rewrites provider URLs. See Adding a New Local Model.

Health checks

curl -s http://localhost:8080/health
{
  "status": "healthy",
  "routes_configured": 3,
  "database_configured": true,
  "database_connected": true,
  "stores": {
    "operational_store": {"status": "ok", "backend": "postgres", "cache": "in_memory"},
    "log_store": {"status": "ok", "backend": "postgres"}
  }
}
  • routes_configured counts published route entries — one per model id in the active registry, plus one per alias. It is whatever your registry defines.

  • database_configured distinguishes “this deployment asked for no database” from “the database is down”: with DB_ENABLED=false it is false and the status is still healthy; with a database configured but unreachable at startup, /health answers 503 with "reason": "database_unavailable_at_startup".

  • status becomes degraded — still HTTP 200 — when one configured store is down but the other is serving. That shape is deliberate: the container HEALTHCHECK uses curl -f /health, so returning 503 for partial degradation would tear down backends that are still answering requests.

Use /health/ready for a strict readiness probe: it applies AND-logic across configured stores and returns 503 unless every one of them is up. /health/deep additionally reports per-endpoint health.

Marking monitor traffic

A monitor that drives real inference — hitting /v1/chat/completions on a schedule to measure a backend end to end, rather than just polling /health — would otherwise land in api_logs, skew the dashboards, and feed RouteWise’s online learning as if it were user demand. Send X-Probe: synthetic on those requests to keep them out: a marked request is left out of api_logs (and its rejections out of the rejection log), does not record a routing observation, and carries an X-Provider response header naming the backend that answered, so the monitor can confirm which route it exercised.

The marker is honoured only from an authenticated internal- or admin-role API key — never a free/pro key, an agent-sandbox grant, or, importantly, an anonymous caller on a deployment running with USER_AUTH_ENABLED=0 (auth-off hands every caller the admin role, which is not the same as holding a monitor identity). From any other caller the header is ignored and the request is logged like ordinary traffic. A deployment that needs probes without auth wants an explicit mechanism — a shared secret, a source allowlist — not this header.

The marker is about noise, not access. It cannot keep a monitor’s own auth failures from tripping the repeated-auth-failure blocklist, because that decision is made before the presented key is read — see A monitor or service account is suddenly getting 429s.

The one thing the marker never touches is billing: cost and quota are incremented unconditionally on every surface, for trusted and untrusted callers alike, so a probe cannot be used to obtain unmetered inference. To keep marked traffic in api_logs after all — to see a monitor’s real latency and spend in the requests dashboard — turn on the log_synthetic_probes runtime setting; the routing-observation and X-Provider behaviour is unchanged.

Alerting

The backend has an in-process alert engine that posts to a Slack webhook. It is off unless you turn it on:

ALERTS_ENABLED=true
SLACK_ALERTS_WEBHOOK_URL=https://hooks.slack.com/services/...

If SLACK_ALERTS_WEBHOOK_URL is empty it falls back to SLACK_WEBHOOK_URL, so one webhook can serve both code paths.

Rules and thresholds are a deployment’s own; this repository ships no alerts file. Point ALERTS_CONFIG_PATH at yours, or leave it unset and the built-in thresholds apply. The rule types and evaluation live in apps/backend/serving/observability/.

Two auth rules are worth knowing apart, because their defaults differ on purpose:

  • rules.auth_failure_spike is off. Bad keys are internet background noise, and a count of them names nothing to act on. The auth_failure log records are emitted regardless.

  • rules.auth_ip_blocked is on. This one fires when the blocklist starts refusing a source — a discrete decision at a much higher threshold, naming an address. It is on because the source is sometimes the deployment’s own; see the troubleshooting entry below.

What an auth alert tells you

An auth-failure card names the sources and, where it can, the accounts behind them — the addresses they came from, how many distinct ones, the leading characters of the keys presented, why each failed, and the paths being hit. The recovery card carries the same picture of the incident that just closed, rather than only the rule’s name: by the time a spike resolves, the window it breached on is empty, so the numbers have to come from a tally kept across the incident.

Two of those lines are worth reading carefully:

  • Known accounts. Most auth failures are anonymous by construction — nobody was authenticated, which is the failure. A named account means a key this deployment did issue was presented and refused, with credential_state saying why (revoked, expired, user_suspended). That is the actionable case: a monitor, CI job or service account whose credential went stale. Resolving the owner costs one indexed lookup per failed auth, bounded by the shared rejection-enrichment budget and shed instantly under a flood; set AUTH_FAILURE_IDENTIFY_CALLER=false to spend nothing and lose the line. Read a named account as evidence and no named account as unknown, never as proof the traffic is external: the lookup is shed during exactly the flood you are investigating, and answers nothing on a timeout, a failed lookup, or with the setting off.

  • Arrived via peers. Present only when the reported addresses did not come off the socket. They are then only as trustworthy as the proxy that set them, and a forged X-Forwarded-For is exactly how a source spreads its failures across the blocklist’s buckets. The line names the sockets they actually arrived on.

Counts marked (capped) are floors, not totals: a source rotating addresses faster than the tally tracks them stops being counted rather than being allowed to grow it without bound. Do not size an incident from a capped number.

Silencing alerts from the dashboard

The admin dashboard’s Settings tab has two controls, and neither changes your alerts file:

  • Slack Alerts snoozes every alert for up to seven days.

  • Alert Types mutes one type of alert, such as auth_ip_blocked or circuit_open, for an hour, a day, a week, or until it is unmuted. Every other type keeps sending.

A mute is stored in the database, so it survives a restart and reaches every gateway process within a few seconds. While a type is muted nothing of it is sent, and an incident that opens during the mute also closes without a recovery message. An incident that was posted before the mute still gets its recovery, and a breach still live when the mute lifts pages at its next evaluation. A mute is for an alert you want back later; to retire a rule for good, turn it off in your alerts file instead.

The first admin account

Every deployment needs one administrator to start with. Choose how you create it with care:

Warning

Do not combine SIGNUP_ENABLED=1, SIGNUP_REQUIRE_EMAIL_VERIFICATION=0 and ADMIN_EMAILS on an instance anyone else can reach. Together they are a privilege-escalation recipe:

  • with verification disabled, POST /auth/signup marks any address as verified without sending mail to it;

  • on login and on every token refresh, the backend promotes any account whose address is listed in ADMIN_EMAILS from free to admin, with no check that the person signing up owns that address.

So a stranger who guesses or reads your ADMIN_EMAILS value signs up with that address and is an admin on their first login.

Pick one of these instead.

Preferred — create the admin out of band and leave ADMIN_EMAILS unset. ops/admin/create_admin.py writes the row directly: it creates the account with role='admin', status='active', email_verified=TRUE, or promotes an existing account with the same address. Run it from the repository root once the backend has started at least once (the backend creates the schema):

python ops/admin/create_admin.py --email [email protected]

Run it inside the project environment (source .venv/bin/activate after make setup-dev) so the serving package is importable. It reads DB_HOST, DB_PORT, DB_NAME, DB_USER and DB_PASSWORD from .env, and Postgres publishes on 127.0.0.1:5432, so it works from the host shell. Omit --password and it prompts, keeping the password out of your shell history. With this in place the instance can run with signup closed:

USER_AUTH_ENABLED=1
SIGNUP_ENABLED=0

Alternative — keep signup open, but leave verification on. SIGNUP_REQUIRE_EMAIL_VERIFICATION defaults to true, and with it on an account cannot log in until it has followed a link sent to the address, which restores the ownership check that ADMIN_EMAILS itself does not perform. This needs working SMTP; without it nobody can complete a signup.

ADMIN_EMAILS also picks the default recipients for signup approval mail. To narrow the notification list without changing who holds the admin role, set SIGNUP_NOTIFY_EMAILS (comma-separated); when it is empty, notifications fall back to ADMIN_EMAILS.

SIGNUP_NOTIFY_EMAILS=[email protected]

Both signup_enabled and signup_require_email_verification can also be flipped at runtime through the settings store, and the runtime value wins over the environment. A .env line is the starting point, not a guarantee.

Database

PostgreSQL 16 runs in the postgres service with its data in the Docker volume hybridinference_postgres_data. A psql shell:

docker exec -it hybridinference-postgres psql -U "${DB_USER}" -d "${DB_NAME}"

Schema details are in Database.

pgAdmin (optional)

make up COMPOSE_PROFILES=admin

pgAdmin then listens on 127.0.0.1:5050 with SCRIPT_NAME=/pgadmin, so an SSH tunnel to that port is enough to reach it. The console can also proxy it at /pgadmin/, gated on an admin session by apps/frontend/src/app/pgadmin/[[...path]]/route.ts, which denies on every unexpected condition, including a backend it cannot reach.

One thing to get right: whether pgAdmin also asks for its own login is set by PGADMIN_CONFIG_SERVER_MODE, and the two defaults disagree. The Compose service falls back to False, which serves pgAdmin with no login at all; .env.example suggests True, which turns pgAdmin’s own login on behind the console’s gate. True is the safer of the two.

Troubleshooting

A service will not start

make logs s=backend
make ps
  • required variable DB_NAME is missing a value: DB_NAME must be set in .env file — Compose stopped at interpolation before starting anything. DB_NAME, DB_USER and DB_PASSWORD are declared required with the ${VAR:?message} form, so the half after the colon is the Compose file’s own text and the most greppable part of the line.

  • Port already in use — override BACKEND_PORT, FRONTEND_PORT or DB_PORT.

  • Database connection failed — check make ps for the postgres health status.

A monitor or service account is suddenly getting 429s

Too many authentication failures from this IP. Temporarily blocked. is the gateway’s own abuse defense, not a provider error and not a quota. Once a source accumulates AUTH_FAILURE_BLOCK_THRESHOLD failed authentications inside AUTH_FAILURE_BLOCK_WINDOW_SEC (200 in a day, by default) it is refused for AUTH_FAILURE_BLOCK_DURATION_SEC (a day).

The awkward case is a caller you own — a status monitor, a CI job, a service account — whose key was rotated, revoked, or never reached its environment. It retries on a schedule, crosses the threshold, and is then refused ahead of the key check, which has two consequences:

  • Repairing the credential does not lift the block. The blocklist is consulted before the presented key is read, so a corrected key gets the same 429 until the deadline passes.

  • The 429 hides the original error. Whatever the caller reports after the block is in place says nothing about whether the underlying 401/403 was fixed.

Which caller is it? Turn on the log_rejected_requests admin setting and the refusals land in Recent Requests as ip_blocked rows. Each row names the account behind the key the caller presented — including a key that was revoked or expired, which is what a stuck monitor is presenting — and labels it with the credential’s state (revoked, expired, user_suspended) beside the user. A row with no user is unresolved, not proof of a stranger: the key may be one this deployment never issued, or the lookup may have been shed, since it runs on a strict budget that gives up first under exactly the flood a block is holding back.

To recover, first fix the credential, then clear the block:

# Which sources is this worker refusing?
curl -s -H "Authorization: Bearer $ADMIN_TOKEN" \
  http://localhost:8080/admin/auth-blocks

# Lift one. `ip` takes a raw address, or a bucket key exactly as listed
# (IPv6 sources are bucketed to their /64).
curl -s -X POST -H "Authorization: Bearer $ADMIN_TOKEN" \
  -H 'Content-Type: application/json' \
  -d '{"ip": "203.0.113.7"}' \
  http://localhost:8080/admin/auth-blocks/clear

cleared: false means there was nothing to lift — it lapsed, or that bucket was never blocked. Clearing grants no immunity: a caller still presenting a bad key is blocked again on crossing the threshold. For a source that should never be blocked at all, list it in AUTH_FAILURE_BLOCK_EXEMPT_IPS (comma-separated addresses or CIDRs) instead.

The blocklist is per-process, in-memory state. On the single-process default both endpoints are exact, and a restart also clears every block. Run multiple workers and each holds its own counts, so a listing shows only the worker that answered and clearing may take more than one call.

Resetting the stack

Stop and start, keeping data:

make down && make up

Destroy the database and start clean. docker compose down -v does not do this. postgres_data is declared external: true in deploy/docker/docker-compose.yml, and Compose never removes an external volume — down -v returns success and leaves it fully intact, so make up comes back on exactly the same data. Remove it by name:

make down
docker volume rm hybridinference_postgres_data
make up   # docker-volumes recreates it empty; Postgres re-initialises

Warning

docker volume rm is irreversible and takes every account, API key and request log with it. Take a pg_dump first if any of it matters.

pgAdmin’s own volume (hybridinference_pgadmin_data) is an ordinary local volume that make down leaves in place. To clear pgAdmin’s saved state, remove it by name as described in Resetting pgAdmin.

Rebuilding after code changes

make build               # all images
make build s=backend     # one service