Grafana looks like a charting tool, and that is where most trouble starts. It stores no metrics. Every panel on every dashboard is a set of live queries, re-run on load, on every refresh and on every zoom, against Prometheus, Loki, a SQL database or anything else with a plugin. A dashboard is therefore a query workload, and the person who builds it decides how expensive and how truthful it is.

This page follows one panel from browser to data source and back, then uses that path to explain the choices that matter: how the query step is chosen, why rate() windows have to follow it, what variables and repeated panels cost, why an alert can disagree with the panel it was copied from, and how to keep dashboards in git. For the product tour of data sources, Loki, Tempo and plugins, read Grafana deep dive; for what to put on a dashboard, read Observability dashboards.

Advertisement

The architecture in one picture

Browserpanel + variablesGrafana server/api/ds/queryData sourceplugin backendPrometheusor Loki, SQLqueriesauth, proxyPromQLData framestime + value columnsframes backTransformationsrun in the browserVisualizationpixelsAlert schedulerserver side, no browsersame query APIExpressionsreduce, math, thresholdSQL database: dashboards, users, alert rulesSQLite by default; MySQL or PostgreSQL for several replicasGrafana stores no metrics.
One panel's round trip. The browser sends queries through the Grafana server, which adds credentials and forwards them through the data source plugin; data frames come back and transformations run in the browser. Alert rules use the same query API from a server-side scheduler and never see browser transformations.

Three consequences follow. The server is mostly stateless: dashboards, users, folders and alert rules live in a SQL database, SQLite by default, so running several replicas means pointing them at a shared MySQL or PostgreSQL database. Data source credentials never reach the browser, because queries go through the server's data source API rather than straight from the page to Prometheus. And anything a panel does after the data arrives, a transformation, an override or a value mapping, exists only in that browser tab.

Max data points, the interval and the step

A time series panel can draw only as many points as it has pixels, so Grafana asks for no more. The panel's max data points defaults to its width in pixels. From the dashboard time range and that number Grafana derives $__interval, roughly range divided by max data points, rounded to a tidy value and never below the min interval. For Prometheus the interval becomes the query step: one evaluation per step, one point per pixel.

Worked example: a seven-day range on a panel about 1,000 pixels wide. Seven days is 604,800 seconds, so the interval lands around ten minutes. Zoom to the last hour and it drops to a few seconds, bounded below by the min interval. The step changes by two orders of magnitude between those views, and every query on the panel has to stay correct at both.

Advertisement

Why rate() windows must follow the step

A rate(x[1m]) evaluated every ten minutes looks at one minute in ten and skips the rest. A five-minute spike between evaluations does not exist on the chart. At the other end, a window shorter than two scrape intervals holds fewer than two samples and rate() returns nothing, which is how panels go blank when someone zooms in.

Grafana's answer is $__rate_interval, defined in its Prometheus documentation as max($__interval + scrape_interval, 4 * scrape_interval). The scrape interval comes from the query's Min step if set, otherwise from the data source's Scrape interval setting, which defaults to 15 seconds. The window always covers the whole step plus one scrape, so no sample is skipped, and never falls below four scrapes, so rate() always has data. The catch is the setting: if Prometheus really scrapes every 60 seconds but the data source still says 15, the floor becomes one minute and zoomed-in panels go empty.

# Request rate per route. $__rate_interval keeps the window wide enough
# for rate() at every zoom level; $route is a multi-value variable.
sum by (route) (
  rate(http_requests_total{job="api", route=~"$route"}[$__rate_interval])
)

# p95 latency from a classic histogram, same window rule.
histogram_quantile(0.95,
  sum by (le, route) (
    rate(http_request_duration_seconds_bucket{job="api", route=~"$route"}[$__rate_interval])
  )
)

# The panel that lies: a fixed [1m] window at a 10m step.
# Prometheus evaluates one 1-minute window every 10 minutes and skips the rest.
sum(rate(http_requests_total{job="api"}[1m]))

Use $__rate_interval for rate(), irate() and increase() in panels, set the data source scrape interval to the real one, and use $__range when you want one number over the whole visible window, such as total errors in a stat panel. The underlying PromQL semantics are covered in Prometheus deep dive.

Variables, repeats and the query multiplier

Template variables make one dashboard serve every service, but each has a price. A query variable runs its own query, such as label_values(http_requests_total, route), when the dashboard loads, and again on every time range change if its refresh is set that way. Chained variables run in sequence, because each depends on the one before. A multi-value variable expands into a regular expression alternation, so selecting All on a high-cardinality label sends a huge matcher to the data source.

Repeats multiply. A row repeated for each of twelve services, with eight panels of two queries each, is up to 192 queries when the whole dashboard is in view. Panels below the fold are lazy-loaded as they scroll into view, unless the server's default_preload setting is on, but a wall display or a tall screen shows everything. Add a 10-second auto refresh and five such screens, and that one page can send about 96 queries per second to Prometheus all day. Grafana allows up to 26 queries per panel, but the real limit is your data source. Collapsed rows are the cheapest fix: panels inside a collapsed row are not queried until the row is expanded.

Measuring a dashboard's cost

You cannot fix what you have not counted. The script below uses the HTTP API with a service account token to list every dashboard, count the queries that run when the dashboard is fully viewed against those deferred in collapsed rows, and show the auto refresh setting. Repeat counts depend on variable values at view time, so the multiplier is an explicit assumption to replace.

import os, requests

GRAFANA = os.environ["GRAFANA_URL"]
H = {"Authorization": f"Bearer {os.environ['GRAFANA_TOKEN']}"}  # service account token

def panels(items):
    for pnl in items:
        yield pnl
        yield from panels(pnl.get("panels", []))   # panels nested in collapsed rows

def cost(dash):
    loaded, deferred = 0, 0
    for pnl in dash.get("panels", []):
        if pnl.get("type") == "row" and pnl.get("collapsed"):
            deferred += sum(len(x.get("targets", [])) for x in panels(pnl["panels"]))
            continue
        if pnl.get("type") == "row":
            continue
        n = len(pnl.get("targets", []))
        if pnl.get("repeat"):
            n *= 10            # assumption: replace with the variable's real option count
        loaded += n
    return loaded, deferred

for hit in requests.get(f"{GRAFANA}/api/search", params={"type": "dash-db"}, headers=H).json():
    dash = requests.get(f"{GRAFANA}/api/dashboards/uid/{hit['uid']}", headers=H).json()["dashboard"]
    loaded, deferred = cost(dash)
    print(f"{loaded:5d} in view  {deferred:5d} deferred  refresh={dash.get('refresh') or 'off':6}  {hit['title']}")

Sort the output by queries in view times refresh rate. A handful of dashboards often account for most of the query load Grafana puts on the data sources. Then fix them in order of impact: move expensive aggregations into recording rules so panels read a precomputed series, collapse rarely used rows, raise min interval on panels that do not need fine resolution, and set a sensible default refresh. The server-wide floor on refresh is min_refresh_interval in the dashboards section of the configuration, 5 seconds by default; raising it stops anyone from saving a 1-second dashboard.

Transformations versus alert expressions

A common surprise: someone builds a panel, adds a transformation to join two queries or filter rows, and then creates an alert from it. The alert fires on different numbers, or never. Transformations run in the browser, after the data arrives. Grafana-managed alert rules run on the server, from a scheduler that evaluates each rule group on its interval, and they use their own processing step, expressions: reduce a series to one number, do math across queries, compare against a threshold.

So the rule is: anything an alert depends on must be in the query or in alert expressions, never in a transformation. Better still, move the logic into the data source, as a recording rule or a SQL view, so panel and alert read the same series.

Alert rule states and the pending period

Each rule produces one alert instance per label set its query returns. An instance is Normal, Pending, Alerting, NoData or Error. A condition that becomes true moves the instance to Pending; only if it stays true for the pending period does it become Alerting and reach a contact point through the notification policy tree, which routes on labels. A pending period of zero alerts on the first true evaluation, which is how flapping pages are born.

NoData and Error are separate decisions. When the query succeeds with no series, or fails outright, you choose per rule whether the instance goes to NoData or Error, to Alerting, to Normal, or keeps its last state. Silence is a signal in its own right: an exporter that died produces no data, and a rule set to Normal on no data hides that. Rules for paging on SLO burn are covered in burn-rate alerting. Many teams keep those rules in the Prometheus or Mimir ruler, next to the data, and use Grafana-managed rules for alerts that combine sources.

Dashboards and data sources as code

Dashboards are JSON documents with a stable uid and a version number. Clicking them together works for one team and fails for fifty: nobody reviews changes, deleted panels are gone, and every environment drifts. Provisioning fixes this. Grafana reads YAML files at startup that declare data sources and dashboard providers, and the file provider reloads dashboard JSON from a directory.

# provisioning/datasources/prometheus.yaml
apiVersion: 1
datasources:
  - name: Prometheus
    type: prometheus
    uid: prom-main            # stable uid, referenced by dashboards
    access: proxy
    url: http://prometheus:9090
    jsonData:
      timeInterval: 30s       # the "Scrape interval" field; match the real scrape
      httpMethod: POST

# provisioning/dashboards/default.yaml
apiVersion: 1
providers:
  - name: team-dashboards
    type: file
    allowUiUpdates: false     # UI edits cannot be saved; change the JSON in git
    options:
      path: /var/lib/grafana/dashboards
      foldersFromFilesStructure: true

Generate the JSON rather than hand-editing it, with a library such as Grafonnet or with the Terraform provider, so twenty service dashboards share one definition. Give every data source a fixed uid and reference it from dashboards; names change, uids should not. Setting allowUiUpdates: false makes git the only way to change a provisioned dashboard, which is the point. Keep a sandbox folder that is not provisioned, so people can still experiment in the UI.

Failure modes

  • Spikes vanish when zooming out: fixed [1m] windows at a wide step. Use $__rate_interval.
  • Panels go blank when zooming in: the data source scrape interval is shorter than the real scrape.
  • Prometheus overloaded at 9 a.m.: repeated rows and short refreshes on shared dashboards; audit and collapse.
  • Alert disagrees with its panel: logic in a transformation; move it to expressions or the data source.
  • Silent outage: NoData mapped to Normal on a rule whose exporter died.
  • Dashboards differ between staging and production: UI edits on unprovisioned dashboards; provision from git.
  • Panel time override ignored: overrides have no effect when the dashboard range is absolute.
  • Variable dropdown takes seconds: a label_values query over a high-cardinality metric; restrict it with a label matcher.

Trade-offs

Grafana's strength is that it owns no data: one pane over every store, no ingestion pipeline to run. The cost is that every view is a live query whose price you set through the step, the windows, the variables and the refresh rate. Fine resolution, big repeats and fast refresh make dashboards feel alive and make data sources hurt. Grafana-managed alerts are flexible across sources but add the server as a dependency of paging; data source rules are closer to the data but limited to one store. Provisioning trades UI convenience for review and reproducibility.

What to do next

  1. Check each Prometheus data source's Scrape interval against the real scrape configuration and fix any mismatch.
  2. Search dashboard JSON for rate( with fixed windows and replace them with $__rate_interval.
  3. Run the audit script and rank dashboards by queries in view times refresh rate.
  4. For the top five, add recording rules, collapse rarely used rows and raise min interval where resolution is not needed.
  5. Set min_refresh_interval to a value your data sources can sustain.
  6. List every Grafana-managed alert whose source panel uses a transformation, and move that logic into expressions.
  7. Review NoData and Error handling on every paging rule.
  8. Provision data sources with fixed uids and move team dashboards into git.
Key takeaway: Grafana stores no data; every panel is a set of live queries whose cost and truthfulness you choose. Max data points and the time range set the step, so rate() windows must follow it through $__rate_interval with a correct scrape interval. Variables, repeats and refresh multiply queries; collapsed rows and recording rules cut them. Alerts run server side and never see transformations. Provision data sources and dashboards from git.