2026 · Ongoing
AI Inference Gateway
An OpenAI-compatible inference gateway exploring request routing, rate limiting, caching, autoscaling, observability, and the operational characteristics of serving AI workloads.
- FastAPI
- Redis
- Kubernetes
- KEDA
- OpenTelemetry
Overview
An experimental serving layer for AI inference workloads.
Most inference projects focus on the model. This one focuses on everything around it: the API surface, the queue, the cache, the scaling signal, and the numbers that tell you whether serving is actually efficient.
What it explores
- API compatibility with existing OpenAI clients
- Authentication and per-key rate limiting
- Request routing across backends
- Semantic caching to avoid duplicate work
- Autoscaling driven by queue depth rather than CPU
- Request queues and backpressure
- Latency and token throughput under load
- Cost visibility per request
Request lifecycle through the gateway.
text
Client
↓
Auth + rate limit
↓
Cache lookup
│
├── hit → response
└── miss ↓
Queue
↓
Inference backend
↓
Metrics + traces