Skip to content

2026 · Ongoing

Production AI Platform

A production-style platform for deploying and operating AI applications on Kubernetes, with infrastructure as code, GitOps delivery, observability, autoscaling, and reliability controls.

  • AWS
  • EKS
  • Terraform
  • Kubernetes
  • Helm
  • Argo CD
  • GitHub Actions
  • Prometheus
  • OpenTelemetry

Overview

Running an API in a container is relatively straightforward.

Operating it reliably introduces a much larger set of problems. How is infrastructure created? How does software reach production? How is traffic routed? What happens when demand suddenly increases? How do we know when something is broken? How do we recover safely?

This project is my attempt to work through those questions as one connected system rather than a collection of unrelated tools.

Architecture

Request path from the public internet to the application and its dependencies.

text
User
│
▼
DNS
│
▼
Load Balancer
│
▼
Ingress
│
▼
Kubernetes
│
├── API
├── Workers
└── Supporting Services
     │
     ├── Redis
     └── PostgreSQL

Infrastructure is provisioned with Terraform. Applications are packaged as containers and deployed through a GitOps workflow. Metrics, logs and traces provide visibility into the system after deployment.

The platform is deliberately boring at the application layer. The interesting work is in everything around it: identity, networking, delivery, scaling, and the signals that tell you whether any of it is working.

Delivery pipeline

From a commit to a reconciled deployment.

text
Developer
 ↓
Git Commit
 ↓
Git Push
 ↓
CI
 ├── Lint
 ├── Test
 ├── Security checks
 └── Build
        ↓
   Container Registry
        ↓
      GitOps
        ↓
     Argo CD
        ↓
   Kubernetes
        ↓
    Verification

The objective was to treat deployment as a system rather than a single kubectl apply. Nothing reaches the cluster without passing through the same path, and the cluster’s desired state is always derived from Git.

A workload is described declaratively, including how it should scale:

yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: inference-api
  namespace: ai-platform
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: inference-api
  minReplicas: 2
  maxReplicas: 20
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 65
  behavior:
    scaleDown:
      stabilizationWindowSeconds: 300

For the parts of the traffic that are more bursty than CPU-driven, the same workload also carries a KEDA ScaledObject driven by queue depth. The two controllers coexist on different signals, which is a decision I revisit often.

Observability

The platform collects three primary signals.

Metrics

Prometheus tracks application and infrastructure behaviour. Examples include request rate, error rate, latency, CPU utilisation, memory utilisation, and pod count.

Logs

Centralised logs allow application events to be investigated without connecting directly to individual containers.

Traces

OpenTelemetry traces follow requests across application boundaries.

Together these signals provide a clearer picture of what failed, where it failed, and what happened before the failure.

SignalValueNote
Signals3metrics · logs · traces
Alert burnMulti-windowfast + slow burn
DashboardsPer-serviceplus platform overview

Failure testing

A system is more interesting when things stop working. I deliberately test scenarios such as:

  • Pod termination
  • Worker termination
  • CPU load and memory pressure
  • Broken dependencies
  • Increased network latency
  • Bad deployments
  • Application crashes

For each failure I examine the same chain:

What I trace for every injected failure.

text
Failure
 ↓
Detection
 ↓
Alert
 ↓
Kubernetes Response
 ↓
Application Recovery
 ↓
User Impact

What I learned

The largest shift was moving from thinking about individual technologies to thinking about the complete request and deployment paths.

Kubernetes is only one component. Terraform, CI/CD, networking, observability, security and application design all interact with it. A platform becomes useful when those pieces behave as one system.

Related writing