Prometheus

Prometheus

Definition: An open-source monitoring system that pulls time-series metrics from configured targets at regular intervals, stores them in its own local time-series database, and exposes a query language for slicing, aggregating, and alerting on that data. Originally built at SoundCloud in 2012, inspired by Google’s internal Borgmon monitoring system, it was released publicly in 2015 and became the second project ever accepted into the Cloud Native Computing Foundation, right after Kubernetes itself. That early pairing with Kubernetes is why Prometheus became the de facto metrics standard for cloud-native infrastructure.

Core Services & Concepts

  • Pull-based scraping — Observability and Monitoring, Prometheus reaches out to targets on a schedule (commonly every 15-60 seconds) rather than waiting for them to push data, which makes it easy to reason about what’s being monitored just by reading the scrape config
  • PromQL — its own functional query language for slicing, aggregating, and forecasting time-series metrics, supporting range queries, rate calculations, and joins across metric labels
  • Exporters — small processes that translate a third-party system’s internal state (a database, a load balancer, a hardware sensor) into the Prometheus text exposition format on an HTTP endpoint
  • Alertmanager — a companion tool that receives firing alerts from Prometheus alert rules and handles deduplication, grouping, silencing, and routing to notification channels (Slack, email, PagerDuty)
  • Labels — key-value pairs attached to every metric (e.g. instance, job, region) that give Prometheus’s data model a multi-dimensional shape, letting PromQL filter and aggregate along any combination of them
  • Service discovery — integrations with Kubernetes, Consul, EC2, and other platforms that let Prometheus automatically discover new scrape targets as infrastructure scales up or down, instead of maintaining a static target list
  • Recording rules — pre-computed, periodically evaluated PromQL expressions saved as new time series, used to keep expensive queries fast on dashboards that are loaded frequently
  • Pushgateway — a bridge component that accepts pushed metrics from short-lived batch jobs, which would otherwise disappear before a scrape could ever reach them
  • Federation — a mechanism for one Prometheus server to scrape aggregated data from another, used to build hierarchical or multi-cluster monitoring topologies

How It Works: The Pull-Based Scraping Model

  • Every target exposes a /metrics HTTP endpoint returning plain-text, line-oriented metric samples; Prometheus itself decides when to fetch it, based on the scrape_interval in its config
  • Each scrape is stored as a new sample in Prometheus’s local time-series database (TSDB), append-only and organized into two-hour blocks on disk for efficient compaction and querying
  • Because Prometheus controls the scrape schedule, it can immediately tell when a target stops responding (up == 0), giving free-of-cost monitoring of scrape health itself
  • PromQL queries operate over this TSDB directly — a query like rate(http_requests_total[5m]) computes a per-second average request rate over a sliding 5-minute window from raw counter samples
  • The pull model deliberately trades off real-time push semantics for operational simplicity: no target needs to know where Prometheus lives, only the reverse, which scales well in dynamic, ephemeral environments
  • Alert rules are evaluated on a separate periodic cycle inside the same Prometheus process, and any rule whose expression returns results is forwarded to Alertmanager as a firing alert

How Pricing Works

  • Prometheus itself is entirely free and open-source (Apache 2.0 license), with no paid tier, license key, or usage limit of any kind for the core server
  • Long-term managed storage is where commercial cost typically enters — hosted options like Grafana Cloud, Amazon Managed Service for Prometheus, or Thanos/Cortex-based platforms charge per sample ingested or per GB stored
  • Self-hosting Prometheus is “free” in licensing terms but not in practice — teams pay in compute, disk, and the engineering time to operate and scale a stateful monitoring system
  • Because it has no built-in high availability or long-term retention, cost usually shows up in the extra tooling (Thanos, Cortex, Mimir) needed to make it production-grade at scale, not in Prometheus itself
  • There is no vendor to negotiate a contract with for the open-source project, contrasting sharply with Datadog’s per-host, per-feature commercial pricing

Pros

  • De facto standard for metrics in the Kubernetes ecosystem, with first-class service discovery and a huge library of ready-made exporters
  • Powerful, purpose-built query language (PromQL) for custom dashboards, alerts, and derived metrics
  • Free and open source with a huge community, no licensing cost or vendor lock-in for the core system
  • Simple operational model — a single static binary plus a text config file, easy to run locally or in a container for evaluation
  • Multi-dimensional label-based data model is far more flexible than flat metric-name-only systems
  • Strong ecosystem integration with Grafana for visualization, forming a well-trodden open-source monitoring pairing

Cons

  • Pull-based model doesn’t fit every architecture well, short-lived batch jobs need the Pushgateway as a workaround, which itself introduces a single point of failure if not run carefully
  • Long-term storage requires additional tooling (Thanos, Cortex, Mimir), Prometheus itself isn’t built for years of retention on a single node
  • No built-in high availability out of the box, running two independent Prometheus servers scraping the same targets is the common workaround, not a native clustering feature
  • PromQL has a real learning curve, particularly around rate/counter semantics and vector matching, that trips up teams new to the tool
  • Local-disk storage means a single Prometheus instance doesn’t scale horizontally without external help, unlike a distributed database
  • Alerting logic and visualization are split across separate tools (Alertmanager, Grafana), requiring more integration work than an all-in-one commercial platform

Comparison: Prometheus vs Grafana vs Datadog

PrometheusGrafanaDatadog
Primary purposeMetrics collection, storage, and alertingVisualization and dashboarding across data sourcesAll-in-one commercial observability (metrics, logs, traces, APM)
Licensing modelOpen-source (Apache 2.0), CNCF graduated projectOpen-source core (AGPL) plus paid Cloud/Enterprise tiersFully commercial SaaS, no free self-hosted tier
Data storageOwn local time-series databaseStores no data itself, queries external data sourcesFully hosted, proprietary backend
Best fitKubernetes-native metrics collection and alert rulesUnifying and visualizing metrics from multiple backendsTeams wanting one managed platform without operating their own stack

Best For

  • Kubernetes-native and cloud-native applications needing real-time metrics collection and alerting
  • Teams comfortable operating open-source infrastructure who want full control and no per-host licensing cost
  • Environments already standardized on Kubernetes service discovery, where Prometheus’s scrape model fits naturally

Real Examples

  • The default metrics backend for most Kubernetes observability stacks, bundled into distributions like the kube-prometheus-stack Helm chart
  • Used internally at large-scale technology companies including SoundCloud (its birthplace), Uber, and DigitalOcean for infrastructure monitoring
  • Cloud providers now ship managed Prometheus-compatible services (Amazon Managed Service for Prometheus, Google Cloud Managed Service for Prometheus), evidence of how standard its query and data model have become

Use Cases

  • Infrastructure and application metrics collection across containers, VMs, and Kubernetes clusters
  • Alerting on service health and performance thresholds via Alertmanager routing
  • Capacity planning and trend analysis using recording rules over historical metric data
  • Monitoring the health of CI/CD pipelines and deployment rollouts alongside tools like GitHub Actions or Jenkins
  • Autoscaling signals — feeding custom metrics into Kubernetes’ Horizontal Pod Autoscaler for Auto-Scaling decisions
  • Black-box monitoring of external endpoints via the Blackbox Exporter, checking HTTP/TCP/ICMP availability from outside a system

Integration Notes & Common Pitfalls

  • Understand counter vs gauge semantics before writing PromQL, applying rate() to a gauge (or forgetting it on a counter) produces silently wrong numbers
  • Cardinality explosions (labels with unbounded values like raw user IDs or full URLs) are the most common cause of Prometheus running out of memory in production
  • Plan for high availability and long-term retention early, bolting on Thanos or Cortex after the fact is more disruptive than designing for it from the start
  • Use relabeling rules carefully during service discovery, misconfigured relabeling is a common source of missing or duplicated targets
  • Pair Prometheus with Grafana rather than relying on its built-in expression browser for anything user-facing, the two are designed to complement each other
  • Set sensible scrape intervals and retention windows up front, over-collecting high-cardinality metrics at short intervals is the fastest way to overwhelm local disk

Code Example

# prometheus.yml — scrape config for a Kubernetes-style target
global:
  scrape_interval: 15s
  evaluation_interval: 15s

scrape_configs:
  - job_name: "web-app"
    static_configs:
      - targets: ["web-app:8080"]
    metrics_path: /metrics

rule_files:
  - "alert_rules.yml"

alerting:
  alertmanagers:
    - static_configs:
        - targets: ["alertmanager:9093"]

Code Example: A PromQL Query and Alert Rule

# Per-second request rate over a 5-minute window, by service
rate(http_requests_total{job="web-app"}[5m])

# Alert rule: fire if error rate exceeds 5% for 10 minutes
- alert: HighErrorRate
  expr: |
    sum(rate(http_requests_total{status=~"5.."}[5m]))
    /
    sum(rate(http_requests_total[5m])) > 0.05
  for: 10m
  labels:
    severity: page
  annotations:
    summary: "Error rate above 5% for {{ $labels.job }}"

Ecosystem

  • Grafana — the standard visualization layer for Prometheus data, see Grafana for its own dedicated entry
  • Thanos / Cortex / Mimir — projects that add horizontal scalability, long-term storage, and global querying on top of vanilla Prometheus
  • Exporters — a large community library covering databases (postgres_exporter), hardware (node_exporter), and cloud services, translating third-party metrics into Prometheus’s format
  • kube-state-metrics — a companion service exposing the state of Kubernetes objects (deployments, pods, nodes) as Prometheus metrics
  • OpenMetrics — the CNCF standard that formalized Prometheus’s exposition format into a vendor-neutral specification other tools can adopt

Best Practices

  • Keep label cardinality bounded, avoid putting raw IDs, timestamps, or full URLs into label values
  • Use recording rules for expensive, frequently-viewed queries instead of recomputing them on every dashboard load
  • Run Alertmanager with sensible grouping and inhibition rules to avoid alert storms during larger outages
  • Version-control scrape configs and alert rules alongside application code, treating monitoring config as part of the deployable system
  • Set retention and storage limits deliberately rather than relying on defaults, and monitor Prometheus’s own resource usage like any other service
  • Adopt a long-term storage solution (Thanos, Cortex, Mimir, or a managed equivalent) before retention needs outgrow a single node, not after

FAQ

Is Prometheus a full observability platform on its own? Not really — it handles metrics well but has no native log or trace storage, teams typically pair it with Grafana for visualization and a separate logging stack for full observability coverage.

Why does Prometheus use a pull model instead of push? The pull model makes it trivial to detect target health (a failed scrape is immediately visible), keeps targets simple since they don’t need to know where Prometheus lives, and fits dynamic environments well through service discovery.

How long does Prometheus retain data by default? 15 days by default on a single node, configurable, though most production setups pair it with a long-term storage solution like Thanos for retention beyond that.

Can Prometheus monitor short-lived batch jobs? Yes, via the Pushgateway, which accepts pushed metrics from jobs that finish before a scheduled scrape could ever reach them.

Does Prometheus replace Datadog? Prometheus plus Grafana together cover much of what Datadog offers for metrics, but Datadog adds logs, APM tracing, and a fully managed experience out of the box that the open-source pairing requires extra tooling to match.

Common Interview Questions

  • “Explain the difference between a counter and a gauge in Prometheus.” — expect an answer distinguishing monotonically increasing counters (requiring rate()) from gauges that can go up or down
  • “How would you handle high cardinality in Prometheus?” — expect discussion of avoiding unbounded label values and monitoring Prometheus’s own memory usage
  • “Why is Prometheus paired with Grafana instead of used alone?” — expect an answer covering Prometheus’s lack of a polished, shareable visualization layer
  • “How does Prometheus achieve high availability?” — expect mention of running duplicate Prometheus servers or adopting Thanos/Cortex/Mimir, since there’s no native clustering

History

  • Built at SoundCloud starting in 2012 by Matt T. Proud and Julius Volz, inspired by Google’s internal Borgmon monitoring system
  • Open-sourced publicly in 2015, quickly gaining traction in the emerging container and microservices ecosystem
  • Became the second project (after Kubernetes) accepted into the Cloud Native Computing Foundation in 2016, cementing its role as the standard cloud-native metrics tool
  • Reached CNCF “graduated” status in 2018, the foundation’s highest maturity tier, alongside Kubernetes
  • The OpenMetrics project formalized Prometheus’s exposition format as a CNCF standard, extending its influence beyond Prometheus itself
  • Continues active development with a large open-source contributor base, alongside a growing ecosystem of managed and horizontally-scalable storage backends

Dig deeper