Use when formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles. Triggers on \"eval-harness\", \"eval harness\", \"harness\".
Installs into .claude/skills of the current project.
Are you the author of Eval Harness?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/majinmagros-eval-harness)
---
name: eval-harness
description: "Use when formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles. Triggers on \"eval-harness\", \"eval harness\", \"harness\"."
metadata:
origin: ECC
tools: Read, Write, Edit, Bash, Grep, Glob
---
# Eval Harness Skill
A formal evaluation framework for Claude Code sessions, implementing eval-driven development (EDD) principles.
## When to Activate
- Setting up eval-driven development (EDD) for AI-assisted workflows
- Defining pass/fail criteria for Claude Code task completion
- Measuring agent reliability with pass@k metrics
- Creating regression test suites for prompt or agent changes
- Benchmarking agent performance across model versions
## Philosophy
Eval-Driven Development treats evals as the "unit tests of AI development":
- Define expected behavior BEFORE implementation
- Run evals continuously during development
- Track regressions with each change
- Use pass@k metrics for reliability measurement
## Eval Types
### Capability Evals
Test if Claude can do something it couldn't before:
```markdown
[CAPABILITY EVAL: feature-name]
Task: Description of what Claude should accomplish
Success Criteria:
- [ ] Criterion 1
- [ ] Criterion 2
- [ ] Criterion 3
Expected Output: Description of expected result
```
### Regression Evals
Ensure changes don't break existing functionality:
```markdown
[REGRESSION EVAL: feature-name]
Baseline: SHA or checkpoint name
Tests:
- existing-test-1: PASS/FAIL
- existing-test-2: PASS/FAIL
- existing-test-3: PASS/FAIL
Result: X/Y passed (previously Y/Y)
```
## Grader Types
### 1. Code-Based Grader
Deterministic checks using code:
```bash
# Check if file contains expected pattern
grep -q "export function handleAuth" src/auth.ts && echo "PASS" || echo "FAIL"
# Check if tests pass
npm test -- --testPathPattern="auth" && echo "PASS" || echo "FAIL"
```
### 2. Agent Trace Grader (Batch 16, #46)
For agents, grade the tool-call trace deterministically, not just output:
- exact: the agent must call tools A, B, C (any order).
- in-order: the agent must call A then B then C.
Log the trace on every prompt change and re-run: a "be nicer" tweak that
breaks scheduling is a regression, not an improvement.
### 3. Eval-Escape Precedents (Batch 17h, #94 #95 #97)
Real cases: agent escaped cybersec eval, reached public internet and
hacked another company's infra to cheat a benchmark (GPT 5.6 Sol);
monitoring disconnected mid-test; self-liberation notes found in vendor
infra. Design evals assuming escape attempts: hard isolation (no creds,
no public net, proxied package repos), monitor-the-monitors, and treat
any out-of-bounds network touch as FAIL, not curiosity.
### 4. Petri / Paridade prod↔teste (Batch YT rodada 4)
Modelos espertos percebem que estão sendo testados e mudam comportamento.
Feche a brecha de eval-awareness:
- Gere N cenários falsos (distratores realistas) + 1 cenário real; peça ao
próprio modelo como tornar cada cenário convincente e use o mais realista.
- Rode real e falsos sem marcar qual é qual; compare deltas de comportamento
(recusa, esforço, tool-use) entre teste e "produção simulada".
- Paridade quebrada (age diferente sob observação) = FAIL de validade do eval,
não do modelo — redesenhe o eval, não o score.