HybridInference
HybridInference is an open-source LLM inference gateway. It puts one OpenAI-compatible HTTP API in front of a mix of inference servers you run yourself — vLLM, SGLang, Ollama, or anything else that speaks the OpenAI API — and hosted provider APIs, then decides per request which of them answers.
The idea the whole system is built around is that the model id a client asks for is decoupled from the endpoint that serves it. One published id can be backed by several routes at once, so traffic across them can be weighted, failed over when an endpoint degrades, priced, and logged — without the client changing a line.
This is the documentation for the gateway software itself: running it, understanding it, extending it, and contributing to it. The source is at github.com/HarvardMadSys/hybridInference, under the MIT license; bugs and questions go to its issue tracker.
What is in the box
An OpenAI-compatible surface —
POST /v1/chat/completions,POST /v1/embeddings,POST /v1/responsesandGET /v1/models, plus an Anthropic Messages surface at/v1/messagesand/anthropic/v1/messagesfor clients that speak that protocol instead.A routing engine — a per-model choice of router, weighted selection among an id’s routes, automatic fallback to the next route, and a per-endpoint circuit breaker that pulls a failing endpoint out of rotation.
Provider adapters — a generic OpenAI-compatible adapter, which also serves your own local servers, alongside dedicated adapters for OpenRouter, Anthropic’s direct API, Claude via Google Vertex, and Gemini.
A web and admin console — a Next.js app for sign-up and login, API keys, usage and recent requests, a chat playground, and an admin area covering users, model visibility, provider keys, route weights and analytics.
Operational storage — Postgres for accounts, API keys and a request log, with the schema created on startup rather than by a migration tool.
The backend is Python (3.10–3.13) under apps/backend/, split into
serving/ (HTTP surface, auth, adapters, storage, observability) and
routing/ (route table, strategies, endpoint health, fallback). The console
is Next.js under apps/frontend/. The repository is MIT-licensed.
Where to start
If you want to |
Start here |
|---|---|
Watch a gateway serve a request, with no account, key or GPU |
|
Run your own gateway against real providers |
|
Choose what your distribution can configure or customize |
|
Follow a request from HTTP through to an upstream call |
|
Serve another model from a provider the gateway already supports |
|
Support a provider whose API the gateway cannot speak yet |
|
Add a provider, a key or a model from the admin console, without editing YAML or restarting |
Configuration, under Runtime configuration from the admin console |
Change how an endpoint gets chosen, or write your own strategy |
|
Send a patch |
|
Identify a version or plan an upgrade |
|
Look up a term such as route, endpoint or distribution |
Scope of this site
These pages document the software, not any one installation of it. Which models a given gateway serves, and how to get an account on it, are the operator’s to publish separately.
That separation is built into the repository: a deployment keeps its identity,
its configuration-file locations and its feature switches in a distribution
overlay under distributions/ rather than in the code, which is why a fresh
clone comes up as nobody’s gateway but your own. Distribution customization
explains how to configure and extend an overlay.
Getting Started
- Quickstart
- Installation
- Configuration
- What the gateway reads at startup
- Where configuration lives
- How a gateway finds its config
- What a fresh clone does with no configuration at all
- The model registry
- The routing file
- Settings that currently have no effect
- Runtime configuration from the admin console
- Environment variables
- Running with your configuration
Concepts
- Architecture
- Routing
- Choosing a router per model
- How
FixedRouterpicks an endpoint - Endpoint health and circuit breaking
- Routes excluded from selection
- Prefill-aware selection
- Session affinity
- Outbound concurrency
- Queue-wait offload
- RouteWise
- Reserving upstream keys for a tier
- API endpoints
- Where the code lives
- Adding a routing strategy
- Routing Internals
- The Public Path Table
- Glossary
Guides
- Adding a New Model
- Adding a New Local Model
- Writing a Provider Adapter
- Step 1: decide whether you need an adapter at all
- Step 2: write the adapter (custom protocols only)
- Step 3: add the model to your registry
- Step 4: configure environment variables
- Step 5: verify it
- Deployment-local adapters
- BaseAdapter API reference
- Advanced features
- Examples
- Troubleshooting
- Best practices
- See also
- Routing through OpenRouter
- Serving a Model from an HPC Cluster
- Claude Code Setup
- Docs RAG Assistant
Customizing a Deployment
Operations
Contributing