Skip to content
Back to skills

Alerting Config

ASecurity

Design effective alerting that catches real issues and minimizes false alarms.

  • 3 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 6, 2026
ai-agentsgobackenddevops

Works with

  • cursor
  • cli

Security analysis

A100/100

Scanned September 6, 2026

npx -y skills add AtulPurohit/Antigravity-Awesome-Skills --skill alerting-config --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Alerting Config?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Alerting Config
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/atulpurohit-alerting-config/badge)](https://www.skillsdirectory.com/skills/atulpurohit-alerting-config)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: alerting-config
description: "Design effective alerting that catches real issues and minimizes false alarms."
category: devops
tags: [alerting,pagerduty,prometheus,monitoring,on-call]
complexity: advanced
risk: low
compatibility: [claude-code, antigravity, gemini-cli, cursor, copilot, codex-cli, autohand, kiro]
source: antigravity-official
version: "1.0.0"
date_added: "2026-07-10"
last_updated: "2026-07-10"
---

# Alerting Configuration Expert

## Purpose
Build alerting systems that notify engineers only for actionable, high-severity events.

## Key Principles
1. Every alert must be actionable
2. Alert on symptoms (high latency) not causes (high CPU)
3. Severity-appropriate routing
4. SLO-based burn rate alerts (better than threshold alerts)

## Prometheus Alert Rules
```yaml
groups:
  - name: slo.rules
    rules:
      - alert: HighErrorBudgetBurn
        # Multi-window, multi-burn-rate alert
        expr: |
          (
            rate(http_requests_total{status=~"5.."}[1h]) /
            rate(http_requests_total[1h]) > 0.14
          ) and (
            rate(http_requests_total{status=~"5.."}[5m]) /
            rate(http_requests_total[5m]) > 0.14
          )
        for: 2m
        labels:
          severity: critical
          team: backend
        annotations:
          summary: "Error rate {{ $value | humanizePercentage }} - burning error budget fast"
          runbook: "https://wiki/runbooks/high-error-rate"

      - alert: HighP99Latency
        expr: histogram_quantile(0.99, rate(http_duration_seconds_bucket[5m])) > 1.0
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "P99 latency is {{ $value | humanizeDuration }}"
```

## AlertManager Routing
```yaml
route:
  group_by: [alertname, service]
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 12h
  receiver: slack-default
  routes:
    - matchers: [severity=critical]
      receiver: pagerduty-critical
      continue: true
    - matchers: [severity=warning]
      receiver: slack-warnings

receivers:
  - name: pagerduty-critical
    pagerduty_configs:
      - service_key: $PAGERDUTY_KEY
  - name: slack-warnings
    slack_configs:
      - channel: '#alerts-warnings'
```

## Outputs
1. Alert rules for application and infrastructure
2. AlertManager routing configuration
3. Runbook template for each alert
4. On-call schedule configuration
5. Alert review process (monthly)

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…