Skip to content

2026 · Ongoing

AI Inference Gateway

An OpenAI-compatible inference gateway exploring request routing, rate limiting, caching, autoscaling, observability, and the operational characteristics of serving AI workloads.

  • FastAPI
  • Redis
  • Kubernetes
  • KEDA
  • OpenTelemetry

Overview

An experimental serving layer for AI inference workloads.

Most inference projects focus on the model. This one focuses on everything around it: the API surface, the queue, the cache, the scaling signal, and the numbers that tell you whether serving is actually efficient.

What it explores

  • API compatibility with existing OpenAI clients
  • Authentication and per-key rate limiting
  • Request routing across backends
  • Semantic caching to avoid duplicate work
  • Autoscaling driven by queue depth rather than CPU
  • Request queues and backpressure
  • Latency and token throughput under load
  • Cost visibility per request

Request lifecycle through the gateway.

text
Client
↓
Auth + rate limit
↓
Cache lookup
│
├── hit  → response
└── miss ↓
      Queue
        ↓
    Inference backend
        ↓
    Metrics + traces

Related writing