Skip to content
Back to skills

Reproduction Experiment Investigator

ASecurity

Specialized investigator for testing hypotheses through reproduction attempts. Designs experiments, executes controlled tests, and reports whether a hypothesis can be confirmed or rejected. Dispatched by debug-conductor.

  • 8 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 6, 2026
ai-agentstestingdebugging

Security analysis

A100/100

Scanned October 6, 2026

npx -y skills add bordenet/superpowers-plus --skill reproduction-experiment-investigator --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Reproduction Experiment Investigator?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Reproduction Experiment Investigator
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/bordenet-reproduction-experiment-investigator/badge)](https://www.skillsdirectory.com/skills/bordenet-reproduction-experiment-investigator)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: reproduction-experiment-investigator
disable-model-invocation: true
source: superpowers-plus
description: "Specialized investigator for testing hypotheses through reproduction attempts. Designs experiments, executes controlled tests, and reports whether a hypothesis can be confirmed or rejected. Dispatched by debug-conductor."
summary: "Use when: testing hypotheses through controlled reproduction attempts."
triggers: []
anti_triggers: []
coordination:
  group: engineering
  order: 10
  requires: ["debug-conductor"]
  enables: []
  escalates_to: ["debug-conductor"]
  internal: true
composition:
  produces: [experiment-evidence, reproduction-recipe, hypothesis-verdict]
  consumes: [hypothesis, reproduction-steps, expected-outcome, environment-context]
  capabilities: [hypothesis-testing, controlled-reproduction, environment-comparison]
  priority: 2
  optional: true
  requires_all: false
---

# Reproduction & Experiment Investigator

> **Role:** Test debugging hypotheses through controlled reproduction attempts. Confirm or reject with evidence.
> **Dispatched by:** `debug-conductor` — never invoked directly by user.
> **Evidence type:** `ExperimentEvidence` (see `skills/_shared/evidence-schema.md`)

## When to Use

Dispatched by `debug-conductor` when a hypothesis needs testing — controlled reproduction, environment comparison, or A/B verification of a suspected root cause.

## Investigation Protocol

### Step 1: Receive and Refine Hypothesis

From the conductor, receive:

- **Hypothesis:** "The failure is caused by [X] when [condition Y] is present"
- **Predicted outcome:** "If hypothesis is correct, we expect [Z] when we [action]"
- **Environment:** Where to reproduce (staging, local, isolated sandbox)

Refine into a testable experiment:

- Define exact steps to reproduce
- Define exact success/failure criteria (not subjective)
- Identify minimum reproduction conditions (strip unnecessary variables)

### Step 2: Environment Assessment

1. Compare reproduction environment to production:
   - Same versions? Same config? Same data? Same load?
   - Document ALL differences — any could explain failure to reproduce
2. If significant differences exist:
   - Can we close the gap? (deploy same version, copy config)
   - If not, document the gap as a limitation

### Step 3: Execute Controlled Reproduction (3+ attempts)

For each attempt:

1. Reset environment to clean state
2. Apply the hypothesized condition
3. Execute the triggering action
4. Record: did the expected failure occur?
5. Record: any unexpected behavior?

**Minimum 3 attempts** — intermittent bugs need statistical confidence.

| Attempts | Reproductions | Confidence |
|----------|---------------|-----------|
| 3/3 | 3 | High (>0.8) — hypothesis strongly supported |
| 2/3 | 2 | Medium (0.5–0.8) — likely correct but intermittent |
| 1/3 | 1 | Low (0.3–0.5) — possible but unreliable |
| 0/3 | 0 | Very Low (<0.3) — hypothesis likely wrong OR environment mismatch |

### Step 4: Control Experiment

If reproduction succeeded:

1. **Remove** the hypothesized condition
2. Re-run the same triggering action
3. If failure disappears → strong confirmation
4. If failure persists → hypothesis may be wrong or incomplete

### Step 5: Produce Evidence

```json
{
  "hypothesis": "Event ordering bug in async pipeline under load",
  "steps": [
    { "action": "Set event processing to async mode", "result": "Config applied", "success": true },
    { "action": "Send 50 concurrent call events", "result": "Events arrived out of order in 12/50 cases", "success": true },
    { "action": "Verify call state machine diverged", "result": "3 calls in disconnected state prematurely", "success": true }
  ],
  "outcome": "reproduced",
  "reproduced": true,
  "attempts": 3,
  "successRate": 1.0,
  "environmentDiff": "Staging uses lower load (50 concurrent vs 500 in prod); reproduction rate may differ"
}
```

## Stop Conditions

- Hypothesis confirmed or rejected with ≥3 attempts
- Reproduction rate established
- Environment assessment shows unbridgeable gap (cannot reproduce here)
- Token budget exhausted
- Wall-clock limit (5 minutes)

## Escalation Conditions

- 0/3 reproduction despite matching environment → hypothesis may be wrong; tell conductor
- Intermittent reproduction (<50%) → flag as timing-dependent; may need more attempts or different conditions
- Environment mismatch prevents reproduction → flag as "cannot test here"; need production access or closer replica

## Common Patterns This Investigator Detects

| Pattern | Evidence Shape |
|---------|---------------|
| **Deterministic bug** | 3/3 reproduction, 0/3 without condition → confirmed |
| **Load-dependent bug** | Reproduces only above certain concurrency threshold |
| **Environment-specific** | Reproduces in prod-like environment but not staging → config/infra difference |
| **Intermittent / race condition** | 1–2/3 reproduction → timing-dependent |
| **Hypothesis disproven** | 0/3 reproduction even with condition → reject hypothesis |

## Failure Modes

| Mode | Symptom | Recovery |
|------|---------|----------|
| Incomplete isolation | Test affected by shared state | Reset environment between experiments |
| False confirmation | Coincidental success in reproduction | Run multiple trials |
| Wrong variable | Testing irrelevant hypothesis | Verify hypothesis matches symptoms |

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…