Skip to content

Experiment

Measuring OpenTelemetry Overhead

An experiment measuring the application cost of increasingly aggressive tracing configurations.

1 min read
  • Observability
  • Performance

The tradeoff

Tracing is not free. Every span is work the application does instead of serving the request. The question is not whether there is a cost, but how the cost scales as sampling and instrumentation get more aggressive.

What I vary

Configuration axis under test.

text
Sampling
  ├── 1%
  ├── 10%
  └── 100%
Instrumentation depth
  ├── HTTP only
  ├── + database
  └── + internal functions
Exporter
  ├── batch
  └── per-request

For each combination I measure throughput, p50/p99 latency, and CPU time, then compare against a baseline with tracing disabled.

What I expect to measure

Head sampling changes the volume of data, not the per-span cost. Instrumentation depth changes the per-request cost. Those two numbers behave very differently, and conflating them leads to the wrong optimisation.

This write-up is being expanded with the full results table and the exact collector configuration.

Related