Skip to content
Back to skills

Chaos

ASecurity

Injects controlled faults for resilience testing on non-prod. Triggers: chaos, fault injection, latency injection, dependency kill, resilience test.

  • 177 stars
  • 0 votes
  • 0 copies
  • 3 views
  • Added May 27, 2026
developmentgobashnodedockerkubernetestestingapibackend

Works with

  • api

Security analysis

A100/100

Scanned May 27, 2026

npx -y skills add softspark/ai-toolkit --skill chaos --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Chaos?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Chaos
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/softspark-chaos/badge)](https://www.skillsdirectory.com/skills/softspark-chaos)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: chaos
description: "Injects controlled faults for resilience testing on non-prod. Triggers: chaos, fault injection, latency injection, dependency kill, resilience test."
effort: medium
disable-model-invocation: true
argument-hint: "[target]"
context: fork
agent: chaos-monkey
allowed-tools: Bash, Read
---

# Chaos Command

$ARGUMENTS

Triggers a controlled resilience experiment.

## Usage

```bash
/chaos <experiment> [target]
# Example: /chaos latency backend-api
# Example: /chaos kill redis
```

## Protocol
1. **Safety Check**: Verify env != PROD.
2. **Baseline**: Check system health is green.
3. **Inject**: Run the fault injection.
4. **Observe**: Monitor logs/metrics for 60s.
5. **Recover**: Restore system health.
6. **Report**: Did we survive?

## Rules

- **MUST** verify target environment is non-production before injecting
- **NEVER** run against a system without a healthy baseline
- **CRITICAL**: abort immediately if recovery does not complete within the observation window
- **MANDATORY**: log every injected fault with timestamp and scope

## Gotchas

- `NODE_ENV=production` on a developer's machine is common — checking that env var alone is not enough proof of non-prod. Combine with kubeconfig context, cloud account ID, or a project-specific env file check before injecting.
- `docker stats` reports cached values; the first sample immediately after injection is often pre-fault. Wait at least 5 seconds before reading metrics.
- Kubernetes liveness probes may self-heal the faulted pod inside the 60s observation window — the report shows green while the workload is still flapping. Check pod restart counters, not just health endpoints.
- Latency injected with `tc` (Linux traffic control) persists across container restarts on the host and across SIGTERM. Always pair the inject step with an explicit `tc qdisc del dev <iface> root` cleanup in the recover step — the `fork` context will not undo it for you.

## When NOT to Use

- In production without an explicit, written runbook — use `/workflow incident-response` for real incidents
- When the system has no observability (no metrics, no logs) — fix observability first
- For load testing — use dedicated load-test tooling, not chaos injection
- During an active incident — stabilize first with `/panic`, then investigate

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…