Skip to content
Back to skills

Observability Slo

ASecurity

Lightweight observability and SLO practice—SLIs, SLO targets, error budgets, logs, metrics, traces, and alerting for apps and data pipelines. Use when defining monitoring, on-call alerts, or reliability targets. Triggers: "SLO", "SLI", "observability", "alerting", "error budget", "monitoring".

  • 25 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added May 27, 2026
data-aiapi

Works with

  • api

Security analysis

A100/100

Scanned May 27, 2026

npx -y skills add charlieviettq/awesome-agent-skill --skill observability-slo --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Observability Slo?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Observability Slo
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/charlieviettq-observability-slo/badge)](https://www.skillsdirectory.com/skills/charlieviettq-observability-slo)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: observability-slo
description: >
  Lightweight observability and SLO practice—SLIs, SLO targets, error budgets,
  logs, metrics, traces, and alerting for apps and data pipelines.
  Use when defining monitoring, on-call alerts, or reliability targets.
  Triggers: "SLO", "SLI", "observability", "alerting", "error budget", "monitoring".
---

# Observability and SLO (lite)

## Three pillars (minimum viable)

| Pillar | Start here |
|--------|------------|
| Logs | Structured JSON; request/job ID; severity |
| Metrics | RED/USE for services; duration, errors, throughput |
| Traces | One trace per request/job across critical hops |

## SLI -> SLO flow

1. **Pick SLI** — measurable user- or business-visible signal.
2. **Set SLO** — target over window (e.g. 99.9% availability / 30d).
3. **Error budget** — 100% - SLO; spend triggers policy (slow features, freeze risky deploys).
4. **Alert on budget burn** — fast burn (pages) vs slow burn (ticket).

## Example SLIs

| Service type | SLI examples |
|--------------|--------------|
| API | Success rate, p95 latency |
| Batch job | Completion within SLA window |
| Model scoring | Scoring latency, feature freshness |

## Alert rules

- Page humans only for SLO-threatening or user-visible outages.
- Warning for degradation trending toward budget burn.
- Every alert links to a **runbook** (symptom, checks, mitigation, escalation).

## Runbook skeleton

```markdown
## Symptom
## Impact
## Checks (ordered)
## Mitigation
## Escalation
```

## Pipeline note

For data/ML jobs: monitor row counts, null spikes, partition lag, and scoring drift alongside infra metrics.

## Anti-patterns

- Alert on every log error without SLO linkage.
- SLOs without measurement (aspirational only).
- Dashboards nobody owns.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…