After the prompts fix, /dm/narrate reached the model call and failed with ConnectError: Connection refused — the container's localhost is not the host, so the default OLLAMA_BASE_URL=http://localhost:11434 hit nothing. Ollama runs on the host (M1's venv smoke reached it because uvicorn ran on the host too; the container never had a route). Add extra_hosts host.docker.internal:host-gateway so the container can reach the host, and document using http://host.docker.internal:11434 in .env.example. Verified a container reaches host Ollama's /api/tags (200). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
/api — FastAPI proxy
The guarding proxy from charter §4. Not a generic API — it exists to keep the API key, the prompts, and the model routing server-side, and to log every call from day one.
Stack: Python, FastAPI. Models: Ollama (dev/staging) → Replicate (prod), routed per-role server-side.
Role endpoints (charter §4)
POST /dm/narrate Narrator — scene/outcome prose (good model)
POST /dm/adjudicate Adjudicator — free text → legal action (small/fast, strict JSON)
POST /dm/improvise Improviser — minor sandboxed event (charter §7 limits)
POST /npc/speak NPC — voice one character (bounded moves, §6)
POST /party/banter Banter — companion callbacks (cacheable, §9)
The client does not know which model serves a role or what the prompt is.
Responsibilities
- Auth · metering (retrofit later — middleware + Stripe webhook, §4)
- Prompt ownership — prompts live in
prompts/, never in the client - Model routing — role → model is config, a deploy not a client patch
- Logging — log seed and full prompt with every call (§10) for replay/eval
Contracts
Every AI response is untrusted (§2). Parse, validate, be ready to discard. One retry on parse failure, then authored fallback (§12).
See docs/ for endpoint schemas, model routing, and deploy notes.