Skip to content
Back to skills

Systematic Debugging

ASecurity

Debug failures to root cause: DNS/TLS, Python CPU/memory with Pyinstrument and Memray.

  • 10 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 23, 2026
ai-agentspythongotestingdebugginggitbackenddocumentation

Security analysis

A100/100

Pro scans all 3 files and shows the line behind each finding

Scanned October 6, 2026

npx -y skills add fmind/dot --skill systematic-debugging --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Systematic Debugging?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Systematic Debugging
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/fmind-systematic-debugging/badge)](https://www.skillsdirectory.com/skills/fmind-systematic-debugging)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: systematic-debugging
description: "Debug failures to root cause: DNS/TLS, Python CPU/memory with Pyinstrument and Memray."
license: MIT
metadata:
  kind: task
  author: Médéric HURIER (Fmind)
  source: github.com/fmind/dot/tree/main/skills/systematic-debugging
  created: "2026-08-08"
  updated: "2026-10-05"
---

# Systematic Debugging

Replace guess-and-check with an evidence loop that localizes where and why behavior diverges; [python-testing](../python-testing/SKILL.md) implements the fix and [incident-response](../incident-response/SKILL.md) owns live outages.

## Workflow

1. **Preserve evidence**: capture the exact error, relevant stack trace, command, inputs, versions, environment differences, timing, and recent changes before touching anything. Keep large sanitized logs in local artifacts; read by incident time, request ID, or error context, retaining the first causal failure and any truncation limits.
1. **Reproduce**: Find the shortest reliable command or sequence; for intermittent failures record the frequency and vary one dimension at a time.
1. **Reduce**: Minimize input, fixture, process count, and component path while keeping the same failure, preferably as a focused test or disposable harness.
1. **Localize**: Trace bad state backward across calls, processes, network boundaries, configuration, and generated artifacts; at each boundary compare what entered with what left.
1. **Find a working comparator**: Locate the nearest known-good test, code path, version, environment, or commit and list every relevant difference before choosing one.
1. **Form one hypothesis**: State `X causes the failure because Y evidence predicts Z observation` and define a minimal probe that could falsify it.
1. **Run the probe**: Change one variable in a reversible fixture or add narrow instrumentation; record whether the prediction held and discard failed hypotheses instead of layering fixes.
1. **Name the root cause**: Explain the triggering condition, the faulty assumption or invariant, the propagation path, and why existing controls missed it; never blame timing, the environment, or a third party until that path and the missing resilience are understood.
1. **Fix within the requested scope**: when the task includes a fix, reuse that authorization, write a regression test for the broken behavior, implement the smallest root-cause correction, and verify the symptom plus affected checks. Broaden to the full gate when repository policy or the change's risk requires it.
1. **Report**: Return symptom and impact, minimal reproduction, evidence and ruled-out hypotheses, root cause and propagation path, the authorized fix or recommended correction, regression proof, and residual uncertainty with the next probe.

## Gotchas

- **Diagnosis does not authorize fixes**: a request to diagnose authorizes investigation, not implementation; observe read-only, reproduce in an isolated temporary directory, and change product code only when the user also asks for a fix.
- **Stop stacking failed fixes**: after three failed fix attempts or hypotheses that expose different shared-state failures, stop stacking fixes and reassess the architecture, reproduction, or problem statement. Summarize what the evidence rules out, continue independent safe probes, and ask only when missing information affects scope, correctness, cost, or reversibility.
- **Instrument every pipeline boundary once**: record presence, shape, identity, status, timestamps, and correlation ids, never secret values; remove the instrumentation unless it has durable value.

## References

- [Dependency resolution](references/dependency-resolution.md): read for package resolver, lockfile, wheel, or build-backend failures before changing any constraint.
- [Network troubleshooting](references/network.md): isolate DNS, connection, TLS, HTTP, proxy, and authentication failures with `doggo`, Python, and `xh` before changing configuration.
- [Python profiling](references/python-profiling.md): read for CPU, allocation growth, or blocked-I/O investigations, including Pyinstrument, Memray, and py-spy captures; use `uv` to run the project Python.

## Documentation

- Upstream: the `superpowers` plugin in `openai/plugins` ships a same-name `systematic-debugging`; preview it, never install it beside this skill ([vendor-skill policy](../agent-project/references/vendor-skills.md#name-collisions)).
- Adapted from [Superpowers systematic-debugging](https://github.com/obra/superpowers/blob/44c9b2d6e889982ac18c27d05a19fefe335194e1/skills/systematic-debugging/SKILL.md), [gstack investigate](https://github.com/garrytan/gstack/blob/960c3a8d6c4d14cb4c5e551a8847f8ec7c4267df/investigate/SKILL.md), [diagnosing-bugs](https://github.com/mattpocock/skills/blob/84fdeffd12f2ee307994d1eb6feb48173b6e0502/skills/engineering/diagnosing-bugs/SKILL.md).
- Companion skills: [repository-history](../repository-history/SKILL.md) (why the code exists).

Files in this skill

  • SKILL.md5.1 KB
  • references/network.md3.7 KB
  • references/python-profiling.md4.2 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…