
The lowest-overhead LLM router. Production-ready, highly available, one OpenAI-compatible endpoint in front of 80 providers and your own vLLM/SGLang — 0.76 µs per request, no I/O on the request path, cache-affinity routing, RBAC, budgets and a 13-screen UI in the binary.