Skip to content
Back to skills

Prometheus

ASecurity

Expert agent for Prometheus 3.x covering PromQL, scrape configuration, service discovery, recording rules, alerting rules, Alertmanager, TSDB management, native histograms, cardinality management, federation, remote write, and Thanos/Mimir integration. WHEN: \"Prometheus\", \"PromQL\", \"scrape config\", \"Alertmanager\", \"recording rule\", \"alerting rule\", \"TSDB\", \"remote write\", \"Thanos\", \"Mimir\", \"Pushgateway\", \"node_exporter\", \"histogram_quantile\", \"rate()\", \"relabel_c...

  • 4 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 24, 2026
devopsgonodeexpresskubernetesapidatabaseperformance

Works with

  • cli
  • api

Security analysis

A100/100

Pro scans all 9 files and shows the line behind each finding

Scanned September 24, 2026

npx -y skills add chrishuffman5/domain-expert --skill prometheus --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Prometheus?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Prometheus
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/chrishuffman5-prometheus/badge)](https://www.skillsdirectory.com/skills/chrishuffman5-prometheus)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: prometheus
description: "Expert agent for Prometheus 3.x covering PromQL, scrape configuration, service discovery, recording rules, alerting rules, Alertmanager, TSDB management, native histograms, cardinality management, federation, remote write, and Thanos/Mimir integration. WHEN: \"Prometheus\", \"PromQL\", \"scrape config\", \"Alertmanager\", \"recording rule\", \"alerting rule\", \"TSDB\", \"remote write\", \"Thanos\", \"Mimir\", \"Pushgateway\", \"node_exporter\", \"histogram_quantile\", \"rate()\", \"relabel_configs\", \"service discovery Prometheus\", \"native histogram\", \"exemplar\", \"cardinality\"."
license: MIT
---

# Prometheus Technology Expert

This skill covers Prometheus 3.x: PromQL, scrape configuration, service discovery, alerting, TSDB internals, and long-term storage integration. Every recommendation addresses the tradeoff triangle: **query performance**, **storage cost**, and **cardinality management**.

## How to Approach Tasks

When you receive a request:

1. **Classify** the request by area:
   - **Architecture** (pull model, TSDB, service discovery, remote write, HA) -- Load `references/architecture.md`
   - **PromQL** (queries, functions, aggregations, binary ops, examples) -- Load `references/promql.md`
   - **Alerting** (alert rules, Alertmanager, routing, receivers, fatigue) -- Load `references/alerting.md`
   - **Configuration** (prometheus.yml, scrape configs, relabeling, recording rules) -- Load `references/configuration.md`
   - **Diagnostics** (TSDB status, cardinality, scrape failures, memory, WAL) -- Load `references/diagnostics.md`

2. **Include cardinality context** -- Never recommend a metric or label without considering its cardinality impact. High cardinality is the primary cause of Prometheus performance issues.

3. **Recommend recording rules** -- For any expensive PromQL expression used in dashboards or alerts, suggest pre-computing it as a recording rule.

4. **Default to native histograms** -- Prometheus 3.x native histograms are more accurate and lower-cardinality than classic histograms. Recommend them for new instrumentation.

5. **Address long-term storage** -- Prometheus local TSDB is designed for short-term retention (15-30 days). For longer retention, recommend remote_write to Thanos, Mimir, or VictoriaMetrics.

## Core Expertise

- **Architecture:** Pull-based scraping, TSDB internals (WAL, blocks, compaction), service discovery (Kubernetes, Consul, EC2, file-based, DNS, HTTP), federation, remote read/write, Pushgateway for short-lived jobs
- **PromQL:** Instant and range vectors, rate/irate/increase, histogram_quantile, aggregation operators (sum/avg/max/min/topk/bottomk), binary operators with vector matching, label manipulation, subqueries
- **Alerting:** Alerting rules (expr, for, labels, annotations), Alertmanager (routing tree, grouping, inhibition, silencing), receivers (Slack, PagerDuty, email, webhook, OpsGenie), Go templating in notifications
- **Configuration:** Scrape configs, relabeling (relabel_configs and metric_relabel_configs), recording rules, OTLP receiver (3.x), global settings, TLS/auth
- **3.x Features:** Native histograms (exponential buckets, server-side quantiles), UTF-8 metric names, OTLP ingestion endpoint, modernized UI

## Prometheus Architecture Quick Reference

```
┌──────────────────────────────────────────────────────────┐
│                    Prometheus Server                      │
│                                                          │
│  ┌──────────┐   ┌──────────┐   ┌──────────────────────┐ │
│  │ Scraper  │   │  TSDB    │   │  HTTP API / PromQL   │ │
│  │ + SD     │──>│  (WAL +  │<──│  + Expression        │ │
│  │          │   │  Blocks) │   │    Browser            │ │
│  └──────────┘   └──────────┘   └──────────────────────┘ │
│       ^                               ^                  │
│       |                               |                  │
│  ┌──────────┐                 ┌──────────────────────┐   │
│  │ Service  │                 │  Rule Manager        │   │
│  │ Discovery│                 │  (Recording + Alert) │   │
│  └──────────┘                 └──────────────────────┘   │
└──────────────────────────────────────────────────────────┘
       ^                               |
       | scrape /metrics               | fire alerts
  ┌────┴─────┐                 ┌───────v────────┐
  │ Targets  │                 │ Alertmanager   │
  └──────────┘                 └────────────────┘
```

**Pull model:** Prometheus scrapes targets at configured intervals. Each target exposes metrics at an HTTP endpoint (default `/metrics`). Targets are discovered dynamically via service discovery or configured statically.

**TSDB:** Custom time-series database. Samples stored in 2-hour blocks. Head block in memory + WAL for crash recovery. XOR compression achieves ~1.3 bytes/sample. Retention by time (default 15d) or size.

## PromQL Quick Reference

### Essential Functions

| Function | Use | Example |
|----------|-----|---------|
| `rate(v[d])` | Per-second rate of counter increase | `rate(http_requests_total[5m])` |
| `increase(v[d])` | Total increase over range | `increase(errors_total[1h])` |
| `histogram_quantile(q, v)` | Percentile from histogram | `histogram_quantile(0.99, sum by (le) (rate(duration_bucket[5m])))` |
| `sum by (l) (v)` | Sum grouped by label | `sum by (job) (rate(requests_total[5m]))` |
| `topk(k, v)` | Top k series by value | `topk(5, sum by (handler) (rate(requests[5m])))` |
| `absent(v)` | Returns 1 if metric missing | `absent(up{job="api"})` |
| `predict_linear(v[d], t)` | Linear prediction | `predict_linear(disk_avail[6h], 4*3600) < 0` |
| `label_replace(v, ...)` | Regex label manipulation | `label_replace(up, "svc", "$1", "instance", "(.+):\\d+")` |

### Common Patterns

**Error rate percentage:**
```promql
100 * sum(rate(http_requests_total{status=~"5.."}[5m])) by (job)
  / sum(rate(http_requests_total[5m])) by (job)
```

**99th percentile latency:**
```promql
histogram_quantile(0.99,
  sum by (le, job) (rate(http_request_duration_seconds_bucket[5m]))
)
```

**CPU utilization per node:**
```promql
1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m]))
```

**Disk fill prediction:**
```promql
predict_linear(node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"}[6h], 4*3600) < 0
```

## Top 10 Operational Rules

1. **Monitor cardinality** -- Check `prometheus_tsdb_head_series` and TSDB status page (`/tsdb-status`) regularly. Each series costs 3-6 KB RAM.

2. **Use recording rules for dashboards** -- Pre-compute expensive `rate()` and `histogram_quantile()` expressions. Name them `level:metric:operations`.

3. **Set `for` on every alert** -- Never fire immediately on transient spikes. Minimum `for: 1m` for critical, `for: 5m` for warnings.

4. **Use `metric_relabel_configs` to drop noise** -- Drop high-cardinality labels and unused metrics at scrape time, not at query time.

5. **Configure remote_write for retention > 30 days** -- Prometheus local storage is not designed for long-term. Use Thanos, Mimir, or VictoriaMetrics.

6. **Run Alertmanager as a cluster (3 nodes)** -- Alertmanager uses gossip for deduplication. Single-node Alertmanager is a SPOF.

7. **Use native histograms (3.x)** -- Exponential buckets eliminate bucket configuration, reduce cardinality, and enable accurate server-side quantiles.

8. **Never put unbounded values in labels** -- User IDs, request IDs, email addresses, and full URLs cause cardinality explosion.

9. **Set `sample_limit` per scrape job** -- Prevent a single target from overwhelming Prometheus with unexpected cardinality.

10. **Use Kubernetes service discovery with relabeling** -- Annotate pods with `prometheus.io/scrape: "true"` and filter in `relabel_configs`.

## Common Pitfalls

**1. Using `irate()` in alerting rules**
`irate()` uses only the last two samples and is too volatile for alerting. Use `rate()` for alerts (smoother, more reliable). Use `irate()` only for dashboards showing momentary spikes.

**2. Forgetting `by (le)` in histogram_quantile**
`histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m]))` fails without grouping by `le`. Always include `le` in the aggregation.

**3. Scraping too frequently**
Scraping every 5s generates 3x more data than 15s. Default 15s is appropriate for most use cases. Only reduce for real-time dashboards on specific targets.

**4. Not setting retention limits**
Default retention is 15 days. Without `--storage.tsdb.retention.size`, disk can fill during traffic spikes. Set both time and size retention.

**5. Using summaries instead of histograms**
Summaries compute quantiles client-side and cannot be aggregated across instances. Histograms allow server-side aggregation. Always prefer histograms.

**6. Treating Prometheus as long-term storage**
Prometheus is optimized for recent data. For queries spanning months or years, use remote_write to a dedicated long-term store (Thanos, Mimir, VictoriaMetrics).

**7. Single Prometheus instance in production**
Run two identical Prometheus instances scraping the same targets. Use Alertmanager cluster for alert deduplication. Use Thanos/Mimir Query for query deduplication.

## Storage Sizing Reference

**Formula:**
```
disk_bytes = (active_series / scrape_interval_seconds) * 1.3 bytes * retention_seconds
```

| Active Series | Scrape Interval | 15-day Retention | 30-day Retention |
|--------------|----------------|-----------------|-----------------|
| 100,000 | 15s | ~7 GB | ~15 GB |
| 500,000 | 15s | ~37 GB | ~75 GB |
| 1,000,000 | 15s | ~75 GB | ~150 GB |
| 5,000,000 | 15s | ~375 GB | ~750 GB |

Add 20% overhead for WAL and compaction scratch space. Use SSD for the WAL directory.

## Reference Files

Load these when you need deep knowledge for a specific area:

- `references/architecture.md` -- Pull model design, TSDB internals (WAL, blocks, compaction, retention), service discovery mechanisms, native histograms, UTF-8 names, OTLP ingestion, federation, remote read/write, Thanos/Mimir integration patterns, HA configurations. Read for architecture and scaling questions.
- `references/promql.md` -- Data types (instant/range vectors, scalars), selectors and matchers, rate functions, histogram functions, aggregation operators, binary operators and vector matching, label manipulation, utility functions, 17+ real-world query examples. Read for any PromQL question.
- `references/alerting.md` -- Alert rule structure and examples, Alertmanager architecture, configuration (routing, grouping, inhibition, silencing), receiver types (Slack, PagerDuty, email, webhook), Go templating, alert fatigue prevention. Read for alerting questions.
- `references/configuration.md` -- prometheus.yml structure, global settings, scrape_configs options, Kubernetes SD scrape config, relabeling patterns (relabel_configs and metric_relabel_configs), recording rules, OTLP receiver config, remote write/read. Read for configuration questions.
- `references/diagnostics.md` -- TSDB status analysis, cardinality investigation, scrape failure diagnosis, high memory troubleshooting, high disk usage, slow query optimization, WAL corruption recovery, Alertmanager diagnostics, key internal metrics. Read for troubleshooting.

## Diagnostic Scripts

Ready-made HTTP API scripts (read-only) in `scripts/`, numbered by investigation order.

- `scripts/01-health-and-targets.sh` -- Down targets with errors, TSDB head stats
- `scripts/02-cardinality-audit.sh` -- Top metrics/label-pairs by series count (the explosion finder)
- `scripts/03-rule-health.sh` -- Broken/slow rules and currently firing alerts

Files in this skill

  • SKILL.md12.2 KB
  • references/alerting.md8 KB
  • references/architecture.md10.9 KB
  • references/configuration.md8.5 KB
  • references/diagnostics.md7.1 KB
  • references/promql.md9.4 KB
  • scripts/01-health-and-targets.sh1.3 KB
  • scripts/02-cardinality-audit.sh1.4 KB
  • scripts/03-rule-health.sh1.4 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…