Skip to content
Back to skills

Grafana Dashboard Builder

ASecurity

Use when building Grafana dashboards backed by Prometheus, Loki, or Tempo, designing PromQL/LogQL queries, wiring template variables, setting alert rules, building SLO dashboards, or maintaining dashboards as code. Triggers: rate() vs increase() confusion, irate vs rate, label_replace, recording rules, alerting rule expressions, multi-dimensional template variables, ad-hoc filters, dashboard JSON model, provisioning via Terraform/grafonnet, p99 / histogram_quantile usage. NOT for Datadog/New ...

  • 2 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 24, 2026
devopsgobashnodeexpressterraformdebuggingapidevops

Works with

  • api

Security analysis

A100/100

Pro scans all 5 files and shows the line behind each finding

Scanned September 24, 2026

npx -y skills add curiositech/port-daddy --skill grafana-dashboard-builder --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Grafana Dashboard Builder?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Grafana Dashboard Builder
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/curiositech-grafana-dashboard-builder-port-daddy/badge)](https://www.skillsdirectory.com/skills/curiositech-grafana-dashboard-builder-port-daddy)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
license: Apache-2.0
name: grafana-dashboard-builder
allowed-tools: Read,Write,Edit,Bash,Glob,Grep,WebSearch,WebFetch
description: 'Use when building Grafana dashboards backed by Prometheus, Loki, or Tempo, designing PromQL/LogQL queries, wiring template variables, setting alert rules, building SLO dashboards, or maintaining dashboards as code. Triggers: rate() vs increase() confusion, irate vs rate, label_replace, recording rules, alerting rule expressions, multi-dimensional template variables, ad-hoc filters, dashboard JSON model, provisioning via Terraform/grafonnet, p99 / histogram_quantile usage. NOT for Datadog/New Relic dashboards (vendor-specific), Grafana plugin development, or Loki ingestion pipeline tuning.'
metadata:
  category: DevOps & Infrastructure
  tags:
    - grafana
    - prometheus
    - promql
    - dashboards
    - slo
    - observability
  provenance:
    kind: first-party
    owners: [port-daddy]
  pairs-with:
    - skill: monitoring-stack-deployer
      reason: Deploys the Prometheus/Loki/Grafana stack these dashboards run on; this skill designs what renders inside it
    - skill: logging-observability
      reason: The producer-side log schema and levels that LogQL panels and error-rate queries depend on
    - skill: observability-apm-expert
      reason: Chooses the metrics, SLIs, and instrumentation that dashboards and alert rules visualize
  io-contract:
    kind: deliverable
    consumes:
      - kind: observability-requirement
        format: markdown
        description: The service, its SLOs, and the questions the dashboard must answer, from a human or another agent.
      - kind: dashboard-plan
        format: json
        description: A structured plan of panel counts, query functions, alert dwell/runbook settings, and provisioning mode, matching schemas/grafana-dashboard-builder-plan.schema.json.
    produces:
      - kind: dashboard-design
        format: markdown
        description: The dashboard layout, PromQL/LogQL queries, recording rules, and alert rules, following this skill's structure and anti-patterns.
      - kind: dashboard-audit-report
        format: json
        description: A deterministic pass/fail audit of the dashboard-plan against this skill's Quality Gates, as produced by scripts/grafana_dashboard_builder_audit.mjs.
---

# Grafana Dashboard Builder

A good dashboard answers one question per panel and one big question per dashboard. PromQL is more expressive than most engineers use; the recurring traps are `rate()` vs `increase()`, label cardinality, and histogram quantile math.

**Jump to your fire:**
- Negative rates / weird step changes → [`rate` vs `increase` vs `irate`](#rate-vs-increase-vs-irate)
- Histogram quantile panel returns NaN → [Histograms and `le`](#histograms-and-le)
- Pager fires from a single noisy blip → [Alert rules](#alert-rules)
- Dashboard has 30 panels nobody reads → [Dashboard structure](#dashboard-structure)
- Same expensive query on many panels → [Recording rules](#recording-rules)
- Multi-value variable returns 5000 series → [Template variables](#template-variables)
- Need to ship a dashboard via PR review → [Dashboards as code](#dashboards-as-code)

## When to use

- New service needs a default dashboard.
- A dashboard exists but is unreadable — too many panels, too many series.
- SLO dashboards (latency p99, error budget burn).
- Alert expressions that fire correctly without paging on transient blips.
- Dashboards-as-code: provisioning via Terraform or grafonnet.

## Core capabilities

### PromQL essentials

```promql
# Per-second request rate over 5min window.
rate(http_requests_total[5m])

# Total requests over 5min.
increase(http_requests_total[5m])

# By status code.
sum by (status) (rate(http_requests_total[5m]))

# Error rate (ratio).
sum(rate(http_requests_total{status=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m]))

# Latency p99 from a histogram metric.
histogram_quantile(0.99,
  sum by (le) (rate(http_request_duration_seconds_bucket[5m]))
)
```

### `rate` vs `increase` vs `irate`

| Function | Returns | Use for |
|----------|---------|---------|
| `rate(m[5m])` | Avg per-second rate over window | Most graphs and alerts. Smooth. |
| `irate(m[5m])` | Instantaneous rate from last 2 samples | Sparkline-style live views. Spiky. |
| `increase(m[5m])` | Total delta over window | "How many requests in 5min." Same as `rate * window_seconds`. |

For alerts, prefer `rate` over `irate` — irate over a noisy counter triggers on every blip.

### Histograms and `le`

Histogram metrics emit `_bucket{le="..."}`, `_sum`, `_count`. To compute quantiles:

```promql
histogram_quantile(0.99,
  sum by (le, route) (rate(http_request_duration_seconds_bucket[5m]))
)
```

Aggregate by `le` AND any dimensions you want to keep in the result. Forgetting `le` returns NaN.

For p99 of *all* requests across routes:

```promql
histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket[5m])))
```

### `label_replace` and renaming

```promql
# Add a `service` label derived from `job`.
label_replace(up, "service", "$1", "job", "(.+)")

# Drop high-cardinality labels for graphing.
sum without (instance, pod) (rate(http_requests_total[5m]))
```

`without` is the cleaner inverse of `by` — sums everything except the listed labels.

### Recording rules

For expensive queries used on many dashboards, pre-compute:

```yaml
# /etc/prometheus/rules/recording.yml
groups:
- name: orders-api
  interval: 30s
  rules:
  - record: job:http_requests:rate5m
    expr: sum by (job, status) (rate(http_requests_total[5m]))
  - record: job:http_request_duration:p99
    expr: histogram_quantile(0.99, sum by (job, le) (rate(http_request_duration_seconds_bucket[5m])))
```

Now dashboards query `job:http_request_duration:p99` instead of recomputing. Cuts dashboard load time and Prometheus CPU.

### Alert rules

```yaml
groups:
- name: orders-api-alerts
  rules:
  - alert: HighErrorRate
    expr: |
      sum(rate(http_requests_total{job="orders-api",status=~"5.."}[5m]))
      /
      sum(rate(http_requests_total{job="orders-api"}[5m]))
      > 0.05
    for: 10m
    labels: { severity: page, team: orders }
    annotations:
      summary: "Error rate >5% on orders-api"
      runbook: "https://runbooks/orders-api/high-error-rate"
```

`for: 10m` is the dwell time — alert only fires after the condition is true for 10 contiguous minutes. Without it, every transient blip pages.

### LogQL (Loki)

```logql
# Last 5 minutes of error logs.
{service="orders-api"} |= "level=error"

# Parse and filter on a JSON field.
{service="orders-api"} | json | status >= 500

# Rate of errors.
sum by (service) (rate({service="orders-api"} |= "level=error" [5m]))

# Latency from a structured field.
{service="orders-api"} | json | unwrap duration_ms | quantile_over_time(0.99, [5m])
```

LogQL extends PromQL with `|=` (contains), `!=`, `|~` (regex), `!~`, `| json`, `| logfmt`, `| unwrap`.

### Template variables

```
Variable: service
Type: query
Query: label_values(up, job)

Variable: instance
Type: query
Query: label_values(up{job="$service"}, instance)
```

`$service` lets the user pick. Multi-value dropdowns + `Include All` cover the common cases.

For ad-hoc filters, use Grafana's "Ad hoc filters" variable type — adds a label filter applied to every panel.

### Dashboard structure

A SLO dashboard typically has:

1. **Top row**: SLO compliance (28d burn, error budget remaining), p99 latency, request rate. Big numbers.
2. **Saturation row**: CPU, memory, queue depth, connection pool usage.
3. **Latency row**: p50/p95/p99, broken down by endpoint or status.
4. **Errors row**: error rate by status, error rate by endpoint, top error messages from logs.
5. **Dependencies row**: downstream service latencies, DB query rate/latency.

Avoid: pie charts (always wrong), gauges with no comparison, tables of 50+ rows. One question per panel.

### Dashboards as code

Terraform:

```hcl
resource "grafana_dashboard" "orders_api" {
  config_json = file("${path.module}/dashboards/orders-api.json")
  folder      = grafana_folder.api.id
  overwrite   = true
}
```

Or grafonnet (Jsonnet) for templated dashboards across services. The point is: dashboards are reviewable, diff-able, and recoverable.

### Annotations

Mark deploys, incidents, and feature flags on graphs:

```promql
# Annotation query — events from a Prometheus metric.
deployment_event{service="orders-api"}
```

Or use Grafana's annotation API to push events from CI.

## Anti-patterns

### `rate()` over a non-counter

**Symptom:** Negative rates, weird step changes.
**Diagnosis:** `rate()` only makes sense on monotonically-increasing counters. Applying to a gauge gives garbage.
**Fix:** `delta()` for gauges, `rate()` for counters. Use the right one.

### `irate` in alerts

**Symptom:** Pager fires from a single noisy blip every few hours.
**Diagnosis:** `irate` reflects the last two samples; one bad sample triggers.
**Fix:** `rate(...)[Nm]` smoothed over minutes; combine with `for: Xm`.

### Histogram quantile without `by (le)`

**Symptom:** Panel shows NaN.
**Diagnosis:** `histogram_quantile` needs the `le` label preserved through aggregation.
**Fix:** `sum by (le, …) (rate(..._bucket[5m]))`.

### Grafana variable that bloats panels

**Symptom:** "Include All" on a 5000-instance variable returns 5000 series.
**Diagnosis:** Multi-value variables with too-broad allowance.
**Fix:** Limit values, scope by another variable, or use `regex` to whittle. Aggregate before display.

### Alert `for:` too short

**Symptom:** Pager fatigue from intermittent network hiccups.
**Diagnosis:** `for: 1m` fires on any blip.
**Fix:** `for: 5m` or `for: 10m` for SLO-tier alerts. Use `for: 0` only for hard-failure metrics ("service down").

### Dashboard with 30 panels

**Symptom:** Slow load, no one reads past the top row.
**Diagnosis:** "Add panel" reflex.
**Fix:** Divide into multiple focused dashboards: SLO, saturation, dependencies, debugging. Cross-link.

## Quality gates

- [ ] Every alert has a `for:` dwell time and a runbook URL.
- [ ] PromQL queries reviewed for `rate` vs `increase` vs `irate` correctness.
- [ ] Histogram quantiles always aggregate by `le`.
- [ ] Recording rules used for queries that appear on >2 panels.
- [ ] Template variables scoped (not multi-select-everything by default).
- [ ] Dashboards stored as JSON in version control; provisioned, not edited live.
- [ ] SLO dashboard has burn-rate and error-budget panels.
- [ ] Annotations for deploys + incidents enabled.
- [ ] Top row of every dashboard answers "is the system healthy" at a glance.

## Deterministic Audit

Before shipping (or reviewing) a dashboard and its alerts, write the plan as JSON matching
`schemas/grafana-dashboard-builder-plan.schema.json` and run it through the deterministic
auditor:

```bash
node scripts/grafana_dashboard_builder_audit.mjs --input examples/sample-input.json
```

`auditGrafanaDashboardBuilder(plan)` (in `scripts/grafana_dashboard_builder_audit.mjs`)
turns this skill's recurring traps into machine-checkable rules over structured fields — no
keyword matching: `rate`/`irate`/`increase` on a gauge (negative-rate garbage),
`histogram_quantile` without `by (le)` (the NaN panel), `irate` inside alert expressions,
SLO alerts with a dwell under 5 minutes, alerts with no runbook URL, expensive queries
shared by >2 panels without a recording rule, panel sprawl past 30, live-edited dashboards
outside version control, and Include-All on a high-cardinality variable. It returns
`{ pass, score, findings, recommendations }` and exits 1 on failure.
`examples/sample-input.json` is a well-formed SLO-dashboard plan (`pass: true`, zero
findings). See `CHANGELOG.md` for the bundle's history.

## NOT for

- **Datadog / New Relic / Honeycomb dashboards** — vendor-specific. No dedicated skill yet.
- **Grafana plugin development** — separate domain. No dedicated skill.
- **Loki ingestion pipeline tuning** — different concern. → `structured-logging-design` for the producer-side schema.
- **Distributed tracing dashboards** (Tempo) — overlapping but distinct. → `opentelemetry-instrumentation` for span/trace generation.
- **Designing the metrics being measured** — this skill assumes metrics exist. → `opentelemetry-instrumentation` for instrumentation patterns.

Files in this skill

  • CHANGELOG.md694 B
  • SKILL.md12.1 KB
  • examples/sample-input.json435 B
  • schemas/grafana-dashboard-builder-plan.schema.json2.9 KB
  • scripts/grafana_dashboard_builder_audit.mjs7.8 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…