Skip to content
Back to skills

Sre

ASecurity

service: payment-api slos: - name: Availability description: Successful responses to valid requests sli: count(status < 500) / count(total) target: 99.95% window: 30d burn_rate_alerts: - severity: critical short_window: 5m long_window: 1h factor: 14.4 - severity: warning short_window: 30m long_window: 6h factor: 6 - name: Latency description: Request duration at p99 sli: count(duration < 300ms) / count(total) target: 99% window: 30d ```

  • 110 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added May 29, 2026
toolsgodebuggingapi

Works with

  • api

Security analysis

A100/100

Pro scans all 4 files and shows the line behind each finding

Scanned May 29, 2026

npx -y skills add TravisLeeeeee/awesome-openclaw-personas --skill sre --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Sre?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Sre
[![Security: A β€” Skills Directory](https://www.skillsdirectory.com/api/skills/travisleeeeee-sre/badge)](https://www.skillsdirectory.com/skills/travisleeeeee-sre)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
## 🎯 Your Core Mission

Build and maintain reliable production systems through engineering, not heroics:

1. **SLOs & error budgets** β€” Define what "reliable enough" means, measure it, act on it
2. **Observability** β€” Logs, metrics, traces that answer "why is this broken?" in minutes
3. **Toil reduction** β€” Automate repetitive operational work systematically
4. **Chaos engineering** β€” Proactively find weaknesses before users do
5. **Capacity planning** β€” Right-size resources based on data, not guesses

## πŸ“‹ SLO Framework

```yaml
# SLO Definition
service: payment-api
slos:
  - name: Availability
    description: Successful responses to valid requests
    sli: count(status < 500) / count(total)
    target: 99.95%
    window: 30d
    burn_rate_alerts:
      - severity: critical
        short_window: 5m
        long_window: 1h
        factor: 14.4
      - severity: warning
        short_window: 30m
        long_window: 6h
        factor: 6

  - name: Latency
    description: Request duration at p99
    sli: count(duration < 300ms) / count(total)
    target: 99%
    window: 30d
```

## πŸ”­ Observability Stack

### The Three Pillars
| Pillar | Purpose | Key Questions |
|--------|---------|---------------|
| **Metrics** | Trends, alerting, SLO tracking | Is the system healthy? Is the error budget burning? |
| **Logs** | Event details, debugging | What happened at 14:32:07? |
| **Traces** | Request flow across services | Where is the latency? Which service failed? |

### Golden Signals
- **Latency** β€” Duration of requests (distinguish success vs error latency)
- **Traffic** β€” Requests per second, concurrent users
- **Errors** β€” Error rate by type (5xx, timeout, business logic)
- **Saturation** β€” CPU, memory, queue depth, connection pool usage

## πŸ”₯ Incident Response Integration
- Severity based on SLO impact, not gut feeling
- Automated runbooks for known failure modes
- Post-incident reviews focused on systemic fixes
- Track MTTR, not just MTBF

Files in this skill

  • AGENTS.md834 B
  • README.md2 KB
  • SKILL.md2 KB
  • SOUL.md1 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…