Docs RAG Assistant
The docs assistant is a retrieval-augmented-generation (RAG) chat feature: it answers users’ questions about a deployment from that deployment’s own public user docs. Both retrieval and generation go through the gateway itself.
It is an early feature with known gaps; see Limitations. To
run it, a deployment needs its documentation as Markdown, an index built from
it, and an API key the assistant can call the gateway with. A checkout with no
distribution has none of these, so the assistant answers 503 until you
supply them.
How it works
/v1/rag/chat is a thin orchestrator: it retrieves in-process, then calls the
gateway’s own public API as a user for the model work.
POST /v1/rag/chat (JWT-gated)
1. embed query ── HTTP ─► POST {RAG_API_BASE_URL}/embeddings (RAG_EMBED_MODEL)
2. cosine top-k over the JSON vector index (in-process)
3. build grounded prompt with citations (in-process)
4. generate ── HTTP ─► POST {RAG_API_BASE_URL}/chat/completions (RAG_CHAT_MODEL)
→ SSE: sources event, then the proxied OpenAI chunks, then [DONE]
Both inner calls present RAG_API_KEY, the assistant’s own API key, so they go
through the normal /v1/embeddings and /v1/chat/completions handlers: they
are logged in api_logs and counted toward cost, quota and concurrency like any
other request. They also carry X-On-Behalf-Of: <user id>, so that usage is
charged to the signed-in user who asked rather than to the shared key, and
counts against that user’s own daily quota. The gateway honours the header only
on requests made with RAG_API_KEY, and ignores it on every other key.
Corpus: the active distribution overlay’s documentation source (
<overlay>/content/docs/docs/source/*.md) — the same markdown that builds that deployment’s public doc site. A checkout with no overlay has no corpus; setRAG_CORPUS_DIRto your own documentation.Vector store: a plain JSON file scanned with pure-Python cosine similarity; a docs corpus is small enough not to need anything more. Like the corpus, the index belongs to the distribution, not to this repository: its default path is
<overlay>/content/rag/docs_index.json, inside whichever distribution the gateway runs.RAG_INDEX_PATHandRAG_CORPUS_DIRoverride these two paths.Why call the gateway as a user (over HTTP) instead of the in-process router? So RAG requests are observable and metered. Calling the router or an adapter directly would skip the per-request logging, cost, quota and concurrency checks that live in the
/v1/*handlers.
Setting it up
Set RAG_API_KEY to a valid user API key, and point RAG_API_BASE_URL at the
gateway’s own address. The default, http://localhost:8080/v1, matches the
port the Compose backend listens on; a deployment that binds the backend
somewhere else must set it, or every RAG request fails at the self-call.
Rebuilding the index
Build the index with the default gateway embedder, which calls a real
embedding model through a gateway:
RAG_CORPUS_DIR=path/to/docs RAG_GATEWAY_API_KEY=hyi-xxx make rag-ingest
RAG_CORPUS_DIR is only needed when your corpus is not the active overlay’s
content/docs/docs/source; without an overlay, ingest fails with a message
telling you to set it.
RAG_GATEWAY_BASE_URL defaults to http://localhost:8080/v1 — embedding is
billable work, so a clone draws on its own gateway rather than on whoever wrote
the default. Point it at the gateway you want to embed through; the key must be
a valid user API key on that gateway. Chunks are embedded with the same
RAG_EMBED_MODEL the serving endpoint uses at query time, so query and
document vectors share one space.
For an offline run with no gateway/key (weak retrieval — dev/CI only):
RAG_EMBEDDER=hash make rag-ingest
The serving endpoint embeds the query with whichever embedder built the index (recorded in the index metadata) and fails loud (HTTP 502) if the query vector’s dimension doesn’t match the index. Always rebuild after switching embedders.
Deploying the index
The index is a deployment artifact, not part of the image build: build it with
make rag-ingest and make the resulting JSON file readable at
RAG_INDEX_PATH. In the Compose deployment the overlay directory is mounted
into the backend, so an index written into the overlay’s content/rag/ is
picked up without an image rebuild, and replacing the file is enough.
/v1/rag/status reports index_loaded, the embedder mode, and chunk count for
a post-deploy check. If the embedding backend is unavailable at query time,
/v1/rag/chat returns a graceful 503 rather than a 500.
Endpoints
Both live under the gateway and authenticate with the dashboard JWT
(get_current_user), so the Next.js chat page calls them with the session token
it already holds.
GET /v1/rag/status— whether the index is built, chunk count, models.POST /v1/rag/chat— body{ messages, top_k?, stream? }. The generation model is fixed server-side (RAG_CHAT_MODEL); it is not client-selectable, so the endpoint can’t be used to reach role-gated models.Streaming (default): SSE — first a
{"type":"sources", ...}event, then OpenAI-format completion chunks, then[DONE].Non-streaming:
{ answer, sources, model }.
A request log row from the assistant has metadata.user_agent = "doc_assistant".
An upstream 429 is passed through to the caller, and may mean the user’s own
daily quota is used up.
curl -sN https://your-gateway.example/v1/rag/chat \
-H "Authorization: Bearer <jwt>" -H 'Content-Type: application/json' \
-d '{"messages":[{"role":"user","content":"How do I get an API key?"}]}'
Turning it off
With an active distribution manifest, features.rag: false turns both
endpoints off for everyone, admins included: they answer 403 before reading
the index or calling a model. true, null or leaving the field out keeps the
assistant on. A broken manifest configuration makes both endpoints answer
503; see Activating a manifest.
The general model APIs are not affected either way.
Configuration
Only RAG_API_KEY is required. The path defaults point inside the active distribution.
Env var |
Default |
Purpose |
|---|---|---|
|
(unset) |
User API key the handler calls the gateway with (required at serving time) |
|
|
Gateway the handler calls (self-call for logging/quota) |
|
the overlay’s |
Vector index location |
|
the overlay’s |
Markdown corpus |
|
|
|
|
|
Gateway used by ingest (gateway mode) |
|
|
Embedding model id (gateway mode) |
|
|
Answer-generation model. Set it to a model your gateway serves: the default names a model this repository does not ship. |
|
falls back to |
User API key ingest presents to |
|
|
Chunks retrieved per query |
|
|
Answer token budget |
|
|
Generation temperature |
Limitations
The index is refreshed manually (
make rag-ingest), not on a schedule — it can lag the docs until regenerated.The
HashEmbedderfallback exists only so the pipeline runs without the gateway (dev/CI); its retrieval quality is weak.No answer caching, no reranking, and history is truncated to the last few turns. The store is loaded into memory per process (cached, mtime-invalidated).
Code map
Path |
Role |
|---|---|
|
Heading-aware markdown chunking |
|
|
|
JSON vector store + cosine search |
|
Prompt assembly + sources payload |
|
|
|
|
|
Chat UI ( |
|
Streaming SSE client |