Grafana
Grafana
Definition: An open-source visualization and dashboarding tool that turns metrics, logs, and traces from dozens of different data sources into readable graphs, tables, and alerts, without storing any of that data itself. Started in 2014 by Torkel Ödegaard as a fork of Kibana focused specifically on time-series graphing, it grew into the standard visualization layer for Prometheus and much of the broader cloud-native monitoring ecosystem. Grafana Labs, the company behind it, later built its own storage backends (Loki for logs, Tempo for traces, Mimir for metrics) to round out a full open-source observability stack around the original dashboarding tool.
Core Services & Concepts
- Data sources — Observability and Monitoring, plugin-based connectors to Prometheus, Loki, Elasticsearch, cloud provider metrics, SQL databases, and dozens of others, queried live rather than imported into Grafana itself
- Dashboards — customizable grids of panels (graphs, tables, gauges, heatmaps) built from queries against one or more connected data sources, savable, versionable, and shareable as JSON
- Panels — the individual visualization units inside a dashboard, each bound to a specific query and visualization type, configurable independently of the rest of the dashboard
- Alerting — unified alert rules that can evaluate queries from any data source and route notifications through contact points (Slack, email, PagerDuty, webhooks), centralizing alerting that used to live in each backend separately
- Variables — dashboard-level placeholders (e.g.
$environment,$namespace) that let a single dashboard template be reused across many services, clusters, or environments by changing a dropdown instead of duplicating dashboards - Explore — an ad hoc, dashboard-free query interface for digging into metrics, logs, or traces during an incident without first building a saved panel
- Provisioning — defining dashboards, data sources, and alert rules as code (YAML/JSON files) so a Grafana instance can be stood up reproducibly instead of clicked together by hand
- Plugins — an extensible system covering additional data sources, panel types, and apps, maintained by both Grafana Labs and the community
How It Works: The Data-Source-Agnostic Query Model
- Grafana holds no time-series data of its own by default — every panel issues a live query to whichever data source it’s bound to, and Grafana’s job is purely to render the result
- Each data source plugin translates Grafana’s generic query editor into that backend’s native query language (PromQL for Prometheus, LogQL for Loki, SQL for relational databases), so users interact with each source’s real syntax rather than a lowest-common-denominator abstraction
- Dashboards are stored as JSON documents describing panels, layout, variables, and data source references, which is what makes them portable, version-controllable, and provisionable from disk
- Alert rules run on a scheduler inside Grafana itself, periodically re-evaluating their underlying query and comparing the result against a threshold, independent of whether the originating data source has its own alerting
- Mixed dashboards can pull from multiple data sources on the same screen — a single row might show Prometheus metrics next to a Loki log panel and a PostgreSQL query, unified by the shared time range control
- Grafana Cloud and Grafana Enterprise add centralized user management, usage-based storage for Loki/Mimir/Tempo, and reporting on top of the same open-source rendering engine
How Pricing Works
- Grafana OSS (the core dashboarding tool) is fully open-source (AGPLv3) and free to self-host indefinitely, with no feature gate on the visualization engine itself
- Grafana Cloud offers a free tier (limited metrics/logs/traces ingestion) plus paid tiers billed on usage — data points ingested, log volume, and active users — rather than a flat per-host fee
- Grafana Enterprise, for self-hosted large organizations, is custom-priced and adds SSO/SAML, enterprise plugins (e.g. Datadog or Splunk data source connectors), and reporting/export features
- Because Grafana itself stores no data, the real cost driver is usually the data sources feeding it (a self-hosted Prometheus, a paid Elasticsearch cluster, or Grafana Cloud’s own hosted Mimir/Loki/Tempo)
- This licensing structure sits in clear contrast to Datadog’s fully commercial, per-host SaaS pricing — Grafana’s core product has no equivalent free ceiling to hit
Pros
- Works with almost any data source, not locked into one metrics backend, which keeps existing monitoring investments usable
- Highly customizable, visually polished dashboards with a large library of panel types and community-built templates
- Strong free/open-source tier, self-hosting the core product costs nothing in licensing
- Centralizes alerting across otherwise-separate backends into one alert rule engine and notification system
- Provisioning-as-code support makes dashboards reproducible and reviewable the same way infrastructure code is
- Large plugin ecosystem and active community, popular dashboards for common stacks are often just an import away
Cons
- Grafana itself doesn’t store metrics, it’s only as good and as fast as the data sources feeding it
- Complex dashboards with many panels and variables can become difficult to maintain and slow to load over time
- Dashboard JSON can drift from what’s provisioned in version control if editors make live changes in the UI without exporting them back
- Alerting configuration migrated significantly between legacy and unified alerting, causing real upgrade friction for teams on older versions
- Managing data source credentials and access control across many connected backends adds its own operational overhead
- Doesn’t collect data itself, so it can’t replace an agent-based platform like Datadog without pairing it with something like Prometheus or Loki
Comparison: Grafana vs Prometheus vs Datadog
| Grafana | Prometheus | Datadog | |
|---|---|---|---|
| Primary purpose | Visualization and dashboarding across data sources | Metrics collection, storage, and alerting | All-in-one commercial observability (metrics, logs, traces, APM) |
| Licensing model | Open-source core (AGPL) plus paid Cloud/Enterprise tiers | Open-source (Apache 2.0), CNCF graduated project | Fully commercial SaaS, no free self-hosted tier |
| Data storage | Stores no data itself, queries external data sources | Own local time-series database | Fully hosted, proprietary backend |
| Best fit | Unifying and visualizing metrics from multiple backends | Kubernetes-native metrics collection and alert rules | Teams wanting one managed platform without operating their own stack |
Best For
- Visualizing metrics from Prometheus, cloud provider metrics, or databases in one unified, shareable view
- Teams already running multiple monitoring backends who need a single pane of glass over all of them
- Organizations wanting dashboard-as-code workflows integrated into existing CI/CD and version control practices
Real Examples
- The standard visualization layer paired with Prometheus across most cloud-native monitoring stacks, commonly bundled together in Helm charts
- Used by large engineering organizations including Bloomberg, PayPal, and eBay for infrastructure and business dashboards
- Grafana Labs’ own hosted product, Grafana Cloud, is used by thousands of companies as a managed alternative to self-hosting the full Prometheus/Loki/Tempo stack
Use Cases
- Infrastructure and application dashboards aggregating metrics from multiple backend systems
- Executive and team-facing status dashboards summarizing system health in business-readable terms
- Cross-referencing metrics, logs, and traces from multiple systems in one screen during incident response
- Capacity planning dashboards tracking long-term resource trends across a fleet of services
- On-call and NOC (network operations center) wall displays showing real-time service health
- Business intelligence–style dashboards built directly on SQL data sources alongside operational metrics
Integration Notes & Common Pitfalls
- Provision dashboards and data sources as code from the start, UI-only dashboards tend to drift and are hard to recover if lost
- Watch dashboard and panel count per screen, too many live queries on one dashboard can overwhelm slower data sources like Elasticsearch
- Be deliberate about who can edit vs. only view dashboards, uncontrolled edit access is a common source of “someone broke the dashboard” incidents
- Migrating from legacy alerting to unified alerting (post-Grafana 8) requires a deliberate review, rules don’t always translate cleanly
- Use template variables consistently across a dashboard so drilling down from an overview to a specific service or namespace stays coherent
- Pair Grafana with a proper data source retention strategy, a beautiful dashboard querying a data source with a 15-day retention window will silently lose historical trends
Code Example
{
"dashboard": {
"title": "Service Overview",
"panels": [
{
"title": "Request Rate",
"type": "timeseries",
"datasource": { "type": "prometheus", "uid": "prometheus" },
"targets": [
{ "expr": "rate(http_requests_total{job=\"web-app\"}[5m])" }
]
}
],
"templating": {
"list": [
{ "name": "environment", "type": "custom", "options": ["prod", "staging"] }
]
}
}
}
Code Example: Provisioning a Data Source as Code
# provisioning/datasources/prometheus.yaml
apiVersion: 1
datasources:
- name: Prometheus
type: prometheus
access: proxy
url: http://prometheus:9090
isDefault: true
jsonData:
timeInterval: "15s"
Ecosystem
- Prometheus — the most common metrics data source paired with Grafana, see Prometheus for its own dedicated entry
- Loki — Grafana Labs’ own log aggregation system, designed to be queried and visualized natively inside Grafana using LogQL
- Tempo — Grafana Labs’ distributed tracing backend, completing the open-source metrics/logs/traces trio alongside Prometheus/Loki
- Grafana Cloud — the managed SaaS version of the full stack, competing directly with platforms like Datadog on convenience while staying open-source underneath
- Community dashboards — a public library of pre-built dashboards (grafana.com/grafana/dashboards) for common data sources and applications, importable by ID
Best Practices
- Store dashboard JSON in version control and provision it, treating dashboards like any other reviewed code change
- Use variables and templating to build one reusable dashboard instead of duplicating near-identical ones per service or environment
- Keep panel counts and query complexity reasonable per dashboard, splitting overloaded dashboards into focused, linked views
- Standardize naming conventions and folder structure early, dashboard sprawl becomes hard to navigate once a team has hundreds of them
- Set meaningful alert thresholds directly tied to user-facing symptoms, not just raw infrastructure metrics, to reduce noisy paging
- Restrict dashboard editing permissions in shared environments and rely on provisioning-as-code for anything that must stay stable
FAQ
Does Grafana store the metrics it displays? No — Grafana is a visualization layer only, it queries external data sources live and stores nothing but its own dashboard definitions and configuration.
Can Grafana replace Prometheus? No, they serve different purposes — Prometheus collects and stores metrics, Grafana visualizes them; most cloud-native stacks run both together rather than choosing one over the other.
Is Grafana only for metrics? No — with Loki and Tempo as data sources it also visualizes logs and distributed traces, making it a general-purpose observability frontend rather than a metrics-only tool.
What’s the difference between Grafana OSS and Grafana Cloud? Grafana OSS is the free, self-hosted core product; Grafana Cloud is Grafana Labs’ managed SaaS offering that also hosts Prometheus-compatible storage (Mimir), Loki, and Tempo so teams don’t have to run any of it themselves.
How does Grafana alerting relate to Prometheus’s own alerting? Grafana’s unified alerting can evaluate queries against any data source, including Prometheus, offering a single alerting system that spans backends instead of configuring Alertmanager and Grafana alerts separately.
Common Interview Questions
- “How does Grafana differ from Prometheus?” — expect a clear distinction between Grafana as a visualization layer and Prometheus as a metrics collection and storage system
- “How would you keep dashboards consistent across environments?” — expect discussion of provisioning-as-code and template variables rather than manually rebuilding dashboards
- “What happens to a Grafana dashboard if its data source goes down?” — expect an answer noting panels will show query errors since Grafana holds no data of its own
- “Why would a team choose Grafana plus Prometheus over Datadog?” — expect a cost/control vs. convenience tradeoff discussion, open-source flexibility against a managed all-in-one platform
History
- Created in 2014 by Torkel Ödegaard as a fork of Kibana, refocused specifically on time-series graphing rather than log search
- Grew quickly alongside Prometheus and the broader Kubernetes ecosystem through the mid-to-late 2010s as the default visualization pairing
- Grafana Labs was founded to commercialize and sustain development of the open-source project, later raising significant venture funding
- Released Loki (2018) and Tempo (2020) to build out its own open-source logs and traces backends, extending beyond pure visualization
- Introduced unified alerting in Grafana 8 (2021), consolidating alert handling across data sources into one engine
- Continues as one of the most widely deployed open-source observability tools, with Grafana Cloud growing as its managed commercial counterpart
Related Terms
Referenced by