Skip to content
Back to skills

Subagent Testing

ASecurity

Test skills via TDD in fresh subagents. Use when validating behavior or preventing bias.

  • 342 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added February 7, 2026
testinggobashrailstesting

Security analysis

A100/100

Pro scans all 2 files and shows the line behind each finding

Scanned September 20, 2026

npx -y skills add athola/claude-night-market --skill subagent-testing --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Subagent Testing?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Subagent Testing
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/athola-subagent-testing/badge)](https://www.skillsdirectory.com/skills/athola-subagent-testing)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: subagent-testing
description: 'Test skills via TDD in fresh subagents. Use when validating behavior or preventing bias.'
alwaysApply: false
category: testing
tags:
- testing
- validation
- TDD
- subagents
- fresh-instances
token_budget: 30
progressive_loading: true
modules:
- modules/testing-patterns.md
model_hint: standard
---
# Subagent Testing - TDD for Skills

Test skills with fresh subagent instances to prevent priming bias and validate effectiveness.

## When NOT To Use

- Writing the skill under test (use `abstract:skill-authoring`)
- A static quality audit with no execution (use `abstract:skills-eval`)

## Overview

**Fresh instances prevent priming:** Each test uses a new Claude conversation to verify
the skill's impact is measured, not conversation history effects.

## Why Fresh Instances Matter

### The Priming Problem
Running tests in the same conversation creates bias:
- Prior context influences responses
- Skill effects get mixed with conversation history
- Can't isolate skill's true impact

### Fresh Instance Benefits
- **Isolation**: Each test starts clean
- **Reproducibility**: Consistent baseline state
- **Measurement**: Clear before/after comparison
- **Validation**: Proves skill effectiveness, not priming

## Testing Methodology

Three-phase TDD-style approach:

### Phase 1: Baseline Testing (RED)
Test without skill to establish baseline behavior.

### Phase 2: With-Skill Testing (GREEN)
Test with skill loaded to measure improvements.

### Phase 3: Rationalization Testing (REFACTOR)
Test skill's anti-rationalization guardrails.

## Quick Start

```bash
# 1. Create baseline tests (without skill)
# Use 5 diverse scenarios
# Document full responses

# 2. Create with-skill tests (fresh instances)
# Load skill explicitly
# Use identical prompts
# Compare to baseline

# 3. Create rationalization tests
# Test anti-rationalization patterns
# Verify guardrails work
```

## Detailed Testing Guide

For complete testing patterns, examples, and templates:
- **[Testing Patterns](modules/testing-patterns.md)** - Full TDD methodology
- **[Test Examples](modules/testing-patterns.md)** - Baseline, with-skill, rationalization tests
- **[Analysis Templates](modules/testing-patterns.md)** - Scoring and comparison frameworks

## Success Criteria

- **Baseline**: Document 5+ diverse baseline scenarios
- **Improvement**: ≥50% improvement in skill-related metrics
- **Consistency**: Results reproducible across fresh instances
- **Rationalization Defense**: Guardrails prevent ≥80% of rationalization attempts

## See Also

- **skill-authoring**: Creating effective skills
- **test-skill**: Automated skill testing command

## Exit Criteria

- [ ] Baseline (RED) phase documents at least 5 diverse scenarios run in fresh Claude instances
  without the skill active, with full response text recorded.
- [ ] With-skill (GREEN) phase uses identical prompts in new fresh instances (not continuations
  of the baseline conversation) and shows >= 50% improvement on skill-related metrics.
- [ ] Rationalization (REFACTOR) phase shows skill guardrails blocking >= 80% of rationalization
  attempts tested across at least 3 pressure scenarios.
- [ ] Results are reproducible: the same prompts in a new fresh instance produce consistent
  outcomes, confirming the effect is not conversation-history priming.

Files in this skill

  • SKILL.md3 KB
  • modules/testing-patterns.md13.7 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…