Observability and Monitoring
Observability and Monitoring
Definition: Observability is the ability to infer the internal state of a complex, distributed system from the external telemetry it emits — commonly described as “the three pillars”: metrics, logs, and traces. Monitoring is closely related but narrower: it’s the practice of watching a predefined set of signals and alerting when they cross known thresholds, answering questions you already thought to ask in advance. The term “observability” itself is borrowed from control theory, where it describes whether a system’s internal state can be determined from its outputs, and was popularized in software engineering roughly around 2016-2018 by engineers at companies like Twitter and Honeycomb who needed to debug distributed systems too complex for dashboards built around anticipated failure modes. In practice the two are used together — monitoring catches the known failure modes, observability is what lets an engineer ask new, unanticipated questions of a system they’ve never had to debug before.
How It Works
- Metrics: numeric, aggregated time-series data — counters, gauges, histograms — collected at regular intervals to measure health and performance cheaply at scale (Prometheus, Datadog, CloudWatch); cheap to store and query but they discard individual-event detail in the aggregation
- Logs: detailed, timestamped text (or structured JSON) events emitted by applications and infrastructure, capturing exactly what happened at a point in time, including error messages, stack traces, and arbitrary context a developer chose to record
- Traces: a record of a single request’s journey across multiple services, built from individual spans (one span per unit of work — an HTTP call, a database query, a function) linked together by a shared trace ID (OpenTelemetry, Jaeger, Zipkin)
- Correlation: the real power of the three pillars shows up when they’re linked — a metric-based alert fires, an engineer pivots to the traces from that time window to find the slow request, then pivots to the specific logs for that trace ID to see the actual error, all via a shared trace/request ID threaded through every layer
- Dashboards and alerting: metrics feed real-time dashboards (Grafana, Datadog) and alerting rules (Prometheus Alertmanager, PagerDuty) that page an on-call engineer when a signal crosses a threshold — this is the “monitoring” half of the discipline, built on top of the same telemetry observability collects
- Instrumentation: none of this happens automatically — applications must be instrumented, either manually (adding log statements, custom metrics, trace spans in code) or via auto-instrumentation libraries (OpenTelemetry SDKs that wrap common frameworks and HTTP clients) that capture the common cases with minimal code changes
Why It Matters
- Critical for rapid debugging of outages, performance bottlenecks, and latency degradation in complex Microservices Architecture systems, where a single user request can touch a dozen services and the failure could be in any of them
- Shortens Mean Time To Resolution (MTTR) during incidents — the difference between “we have no idea why this is slow” and “here’s the exact span that’s taking 4 seconds” is often the difference between a 5-minute fix and a multi-hour outage
- Enables proactive detection of problems before they become customer-facing outages, a slowly climbing p99 latency or error rate is a warning sign monitoring can catch long before a full failure
- Provides the evidence base for capacity planning, SLA reporting, and post-incident reviews — without recorded telemetry, “how often does this actually happen” is a guess instead of a query
Under the Hood: The Three Pillars and How a Trace Gets Assembled
Metrics, logs, and traces solve different problems and trade off differently on cost, cardinality, and detail: metrics are pre-aggregated so they’re cheap to store and fast to query even at massive scale, but they can’t tell you about one specific request, only trends across many; logs preserve full per-event detail but are expensive to store and search at high volume, especially with high-cardinality fields like user IDs; traces sit in between, sampling a subset of requests to preserve the full causal structure of specific transactions without the cost of tracing every single one. A distributed trace is assembled from spans rather than recorded as a single unit: when a request enters the system, the first service generates a trace ID and a root span, then propagates that trace ID (via HTTP headers like traceparent in the W3C Trace Context standard) to every downstream service it calls; each of those services creates its own child span, tagged with the same trace ID and a reference to its parent span ID, records its own start time, duration, and any errors, and propagates the trace ID further if it calls anything else. None of these spans need to be created or transmitted in the same process, or even reach the tracing backend in order — they’re collected independently (often via a local agent or sidecar that batches and forwards them) and only stitched back together into the familiar waterfall view when the backend receives all the spans sharing a trace ID and reconstructs the parent-child tree from the span IDs. This is why a trace can show you that a 2-second request spent 1.8 seconds inside one specific downstream database call, three services deep, something no amount of staring at aggregate metrics or grepping logs service-by-service would surface nearly as fast.
Sampling decisions matter as much as the collection mechanism itself: head-based sampling decides whether to keep a trace at the moment the root span starts, before anything is known about how the request will turn out, which is cheap but risks discarding exactly the slow or erroring requests engineers care about most; tail-based sampling buffers all spans for a trace until it completes and only then decides whether to keep it, which lets a system guarantee “always keep traces with errors or high latency” at the cost of holding spans in memory until the full trace resolves. Most production tracing setups land on a hybrid — a low uniform sampling rate for normal traffic to keep baseline cost down, combined with tail-based rules that always capture the outliers that actually matter during an incident.
Comparison: Observability vs Monitoring-Only vs Logging-Only
| Full Observability (Metrics + Logs + Traces) | Monitoring-Only | Logging-Only | |
|---|---|---|---|
| Question it answers | “Why is this specific request slow?” — even ones you didn’t anticipate | “Is the known signal within a healthy range?” | “What exactly happened at this point in time?” |
| Debugging unknown failure modes | Strong — traces reveal causal paths you didn’t predict | Weak — only surfaces what was pre-defined as a signal | Moderate — detail exists but correlating across services is manual and slow |
| Cost/storage | Highest — three data types, plus correlation tooling | Lowest — small number of aggregated series | High at scale — full-text detail on every event |
| Best fit | Complex distributed/microservice systems | Simple systems with well-understood failure modes | Debugging a single service or low-traffic system |
Common Pitfalls
- Relying solely on system metrics without distributed tracing when debugging microservice latency spikes — metrics can tell you that p99 latency jumped, but not which downstream call in which specific request caused it
- Unstructured, free-text logs that are cheap to write but expensive to query at incident time, structured (JSON) logging with consistent fields pays for itself the first time someone needs to filter by user ID or trace ID under pressure
- Alert fatigue from too many low-value alerts, or thresholds set arbitrarily rather than tied to actual user-facing impact, teams start ignoring pages, which defeats the entire point of monitoring
- Tracing every request in a high-throughput system without sampling, generating enormous data volume and cost for marginal debugging value beyond what a representative sample already provides
- Instrumenting a service in isolation without propagating trace context to downstream calls, breaking the trace into disconnected fragments instead of one continuous picture of the request
- Treating dashboards as the end goal rather than the entry point — a dashboard that shows something is wrong is only useful if it leads somewhere (logs, traces) that shows why
Code Example
# Structured logging + OpenTelemetry span, correlated via trace context
import logging, json
from opentelemetry import trace
tracer = trace.get_tracer("orders-service")
logger = logging.getLogger("orders")
def process_order(order_id: str):
with tracer.start_as_current_span("process_order") as span:
span.set_attribute("order.id", order_id)
trace_id = format(span.get_span_context().trace_id, "032x")
logger.info(json.dumps({
"event": "order_processing_started",
"order_id": order_id,
"trace_id": trace_id, # ties this log line back to the full trace
}))
try:
charge_payment(order_id) # this call propagates trace context downstream
except PaymentError as e:
span.record_exception(e)
span.set_status(trace.StatusCode.ERROR)
raise
A matching Prometheus alert rule that pages when the error rate itself crosses a threshold, giving the metric that would trigger the pivot into the trace above:
groups:
- name: orders-service-alerts
rules:
- alert: HighOrderErrorRate
expr: |
rate(orders_processed_total{status="error"}[5m])
/ rate(orders_processed_total[5m]) > 0.02
for: 5m
labels:
severity: page
annotations:
summary: "Order error rate above 2% for 5 minutes"
Best Practices
- Instrument early, not after the first unexplainable incident — retrofitting tracing onto a mature distributed system under incident pressure is far harder than building it in from the start
- Use structured logging with consistent field names (trace ID, user ID, service name) across every service, so logs are queryable and correlatable rather than just readable one at a time
- Sample traces intelligently rather than uniformly — capture 100% of errors and slow requests, and a smaller representative percentage of normal fast requests, to control cost without losing the signal that matters most
- Set alert thresholds based on user-facing impact (error rate, latency SLOs) rather than arbitrary system internals, and prune alerts that never lead to action
- Adopt OpenTelemetry for instrumentation where possible, a vendor-neutral standard avoids locking telemetry format to one backend and makes switching tools later far less painful
FAQ
Is observability just monitoring with a new name? No — monitoring watches predefined signals for known failure modes, observability is the broader capability to explore and debug unknown failure modes after the fact, using rich telemetry rather than pre-built dashboards alone.
Do I need all three pillars, or can I start with just one? Most teams start with metrics and logs since they’re cheaper and simpler to set up, and add distributed tracing once the system has enough services that “which one is slow” stops being answerable by inspection alone.
What is a “span” exactly, and how is it different from a log line? A span represents a timed unit of work with a start, duration, and parent-child relationship to other spans in the same trace, where a log line is just a timestamped message — spans exist specifically to reconstruct causal, cross-service request flow, which individual log lines don’t capture on their own.
History
- The term “observability” originates in 1960s control theory (Rudolf Kálmán), describing whether a system’s internal state is inferable from its external outputs — software engineering borrowed the word, not the math, decades later
- Traditional monitoring tools (Nagios in the late 1990s, later Graphite and early Prometheus-style time-series systems) established the metrics-and-alerting foundation long before “observability” became a distinct term in software circles
- Distributed tracing was popularized by Google’s 2010 Dapper paper, which described how Google traced requests across its own massive internal service graph, directly inspiring open-source tracing systems like Zipkin and Jaeger
- The OpenTelemetry project formed in 2019 from the merger of OpenTracing and OpenCensus, consolidating instrumentation standards under the CNCF and making vendor-neutral, portable telemetry the default expectation for new tooling
Related Terms
- Container Orchestration and Kubernetes
- Kubernetes (K8s)
- Microservices Architecture
- Auto-Scaling
- Service Mesh
- API Gateway
Example
Prometheus scrapes CPU and memory metrics from every service every 15 seconds, feeding a Grafana dashboard that pages an on-call engineer when error rate crosses 2%. The engineer pivots from that alert into Jaeger, finds the slow trace, and follows its trace ID into the structured logs for the one downstream service actually responsible — going from “something is wrong” to “this specific database query in this specific service is timing out” in minutes instead of hours.
Referenced by