When a customer says "checkout was slow yesterday at 3 p.m.", how long does it take your team to find out why? With logs scattered across services and metrics in a separate dashboard, the answer is often "hours, if ever". Observability is the ability to ask new questions about a running system without shipping new code. OpenTelemetry has become the standard way to get there: one open specification and set of SDKs for traces, metrics and logs, supported by virtually every monitoring vendor and open-source backend. This guide covers how it works, how to instrument Java and Node.js services, and how to keep the data useful and affordable.
Why OpenTelemetry
Before OpenTelemetry, each monitoring vendor had its own agents and SDKs. Switching vendors meant re-instrumenting everything. OpenTelemetry, a CNCF project formed from the merger of OpenTracing and OpenCensus, separates instrumentation from backends (CNCF: OpenTelemetry; OpenTelemetry docs):
- You instrument once with OpenTelemetry APIs, SDKs and auto-instrumentation.
- Data travels in the standard OTLP protocol (OTLP specification).
- You send it to any backend: open-source stacks such as Grafana Tempo, Prometheus or VictoriaMetrics, and Loki, or commercial platforms — and can change your mind later.
The three signals (and a fourth)
| Signal | Answers | Example |
|---|---|---|
| Traces | Where did this request spend its time, across services? | Checkout → payment API took 2.1 s of 2.4 s |
| Metrics | How is the system behaving overall? | p95 latency, error rate, queue depth |
| Logs | What exactly happened at this point? | "Payment declined: insufficient funds" |
| Profiles | Which code consumed CPU or memory? | Newest signal, still in development |
The power comes from correlation: a log line carries the trace ID, so you jump from a slow trace to the exact logs of that request, and from a metric spike to example traces (OpenTelemetry: signals).
Maturity varies by language. In Java, traces, metrics and logs are all stable; in JavaScript and Python, traces and metrics are stable while logs are still in development; profiles are in development (OpenTelemetry status). Check the status for your stack before relying on a signal.
Architecture: SDKs, Collector, backends
[Service A] ─┐
[Service B] ─┼── OTLP ──> [OpenTelemetry Collector] ──> traces → Tempo / Jaeger / vendor
[Service C] ─┘ (receive, process, export) ──> metrics → Prometheus / VictoriaMetrics
──> logs → Loki / OpenSearch
The Collector is a vendor-neutral proxy that receives telemetry, processes it — batching, filtering, sampling, removing sensitive attributes — and exports it to one or more backends (OpenTelemetry Collector). Run it as an agent next to services (a DaemonSet in Kubernetes) and optionally as a central gateway. Changing backends then means changing Collector configuration, not application code.
# otel-collector.yaml (excerpt)
receivers:
otlp:
protocols: { grpc: {}, http: {} }
processors:
batch: {}
attributes/scrub:
actions:
- key: http.request.header.authorization
action: delete
tail_sampling:
policies:
- { name: errors, type: status_code, status_code: { status_codes: [ERROR] } }
- { name: slow, type: latency, latency: { threshold_ms: 1000 } }
- { name: baseline, type: probabilistic, probabilistic: { sampling_percentage: 10 } }
exporters:
otlphttp/tempo: { endpoint: http://tempo:4318 }
prometheusremotewrite: { endpoint: http://victoriametrics:8428/api/v1/write }
service:
pipelines:
traces: { receivers: [otlp], processors: [attributes/scrub, tail_sampling, batch], exporters: [otlphttp/tempo] }
metrics: { receivers: [otlp], processors: [batch], exporters: [prometheusremotewrite] }
Instrumenting services
Java and Spring Boot
The OpenTelemetry Java agent instruments common libraries — HTTP servers and clients, JDBC, Kafka, gRPC, Redis — without code changes (Java agent):
java -javaagent:opentelemetry-javaagent.jar \
-Dotel.service.name=orders-service \
-Dotel.exporter.otlp.endpoint=http://otel-collector:4318 \
-jar orders-service.jar
Spring Boot 4 also offers a dedicated spring-boot-starter-opentelemetry for applications that prefer configuration over an agent, as described in Spring Boot + Kotlin in 2026. With virtual threads, traces are the quickest way to see where requests actually wait, see Virtual threads in Spring Boot.
Node.js and Next.js
The Node.js SDK with auto-instrumentations covers HTTP, Express, popular database clients and more (OpenTelemetry JavaScript). In Next.js, register it in instrumentation.ts so it loads before your application code.
Custom spans and attributes
Auto-instrumentation shows the plumbing; business context makes traces useful. Add spans around important operations and attributes such as tenant, order value or feature flag, following the semantic conventions for standard names (semantic conventions). Never put personal data or secrets in attributes.
Context propagation
Traces work across services because each outgoing request carries the trace context in the W3C traceparent header (W3C Trace Context). Make sure gateways, proxies and message brokers propagate it — Kafka headers, for example — or traces break into disconnected pieces.
Sampling and cost control
Telemetry volume grows fast, and so does the bill. Control it deliberately (OpenTelemetry: sampling):
- Head sampling decides at the start of a trace (keep 10%). Cheap, but may drop the interesting traces.
- Tail sampling in the Collector decides after the trace completes: keep all errors and slow requests, plus a baseline percentage of normal traffic.
- Metrics over logs for things you count. A counter is far cheaper than a log line per event.
- Watch cardinality. Metric labels with unbounded values — user IDs, URLs with IDs — explode storage.
- Retention tiers. Keep detailed traces for days, aggregated metrics for months.
What to measure first
Start with the signals that tell you whether users are happy, then drill down:
- RED for services: request rate, errors and duration for every endpoint.
- USE for resources: utilization, saturation and errors for CPU, memory, disk, connection pools.
- Business metrics: orders per minute, payment success rate, sign-ups.
- SLOs and alerts on symptoms users feel — error rate, latency — rather than every CPU spike, as Google's SRE book recommends (Google SRE: monitoring distributed systems).
Databases and backups deserve the same attention; see PostgreSQL backup and disaster recovery. In a modular monolith, traces also reveal calls between modules long before you split anything out, see Modular monolith with Spring Modulith. On self-managed Kubernetes, we run this stack with VictoriaMetrics, Grafana and Loki, as described in Kubernetes on Hetzner.
Rollout plan
- Deploy the Collector and one backend per signal.
- Add auto-instrumentation to two or three critical services; verify traces connect end to end.
- Add trace IDs to logs and switch logging to structured JSON.
- Define RED dashboards and two or three SLO-based alerts.
- Add custom spans for key business operations.
- Introduce tail sampling and cardinality limits before volume grows.
- Roll out to remaining services with a shared configuration library.
Mistakes to avoid
- Instrumenting everything but alerting on nothing that users feel.
- Logging request bodies with personal data into telemetry backends.
- Using user IDs or full URLs as metric labels.
- Running a different SDK configuration in every service instead of a shared library.
- Keeping 100% of traces forever, then cutting observability when the bill arrives.
FAQ
Does OpenTelemetry replace Prometheus or Grafana? No. OpenTelemetry handles instrumentation and transport; Prometheus, Grafana and others store and visualize the data.
Is auto-instrumentation enough? It is a great start. Add custom spans and attributes for business context, or traces show only technical plumbing.
How much overhead does it add? Typically small for well-configured SDKs with batching and sampling. Measure on your own services under load.
Can we use OpenTelemetry with a commercial vendor? Yes. Most vendors accept OTLP directly or via the Collector, which keeps you free to switch later.
Do we need the Collector for a small setup? You can export directly from SDKs to a backend, but the Collector makes it easy to scrub sensitive data, sample and switch backends later, so we add it from the start.
Sources
- OpenTelemetry. Documentation, Signals, Status.
- OpenTelemetry. Collector, OTLP, Sampling, Semantic conventions.
- OpenTelemetry. Java agent and JavaScript.
- W3C. Trace Context.
- Google. SRE book: Monitoring distributed systems.
- CNCF. OpenTelemetry project.