Skip to content
Back to skills

Eval Harness

ASecurity

Use when formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles. Triggers on \"eval-harness\", \"eval harness\", \"harness\".

  • 2 stars
  • 0 votes
  • 0 copies
  • 3 views
  • Added September 19, 2026
ai-agentsbashperformance

Works with

  • claude code

Security analysis

A100/100

Pro scans all 5 files and shows the line behind each finding

Scanned September 19, 2026

npx -y skills add majinmagros/magros.ai-skills --skill eval-harness --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Eval Harness?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Eval Harness
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/majinmagros-eval-harness/badge)](https://www.skillsdirectory.com/skills/majinmagros-eval-harness)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: eval-harness
description: "Use when formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles. Triggers on \"eval-harness\", \"eval harness\", \"harness\"."
metadata:
  origin: ECC
tools: Read, Write, Edit, Bash, Grep, Glob
---

# Eval Harness Skill

A formal evaluation framework for Claude Code sessions, implementing eval-driven development (EDD) principles.

## When to Activate

- Setting up eval-driven development (EDD) for AI-assisted workflows
- Defining pass/fail criteria for Claude Code task completion
- Measuring agent reliability with pass@k metrics
- Creating regression test suites for prompt or agent changes
- Benchmarking agent performance across model versions

## Philosophy

Eval-Driven Development treats evals as the "unit tests of AI development":
- Define expected behavior BEFORE implementation
- Run evals continuously during development
- Track regressions with each change
- Use pass@k metrics for reliability measurement

## Eval Types

### Capability Evals
Test if Claude can do something it couldn't before:
```markdown
[CAPABILITY EVAL: feature-name]
Task: Description of what Claude should accomplish
Success Criteria:
  - [ ] Criterion 1
  - [ ] Criterion 2
  - [ ] Criterion 3
Expected Output: Description of expected result
```

### Regression Evals
Ensure changes don't break existing functionality:
```markdown
[REGRESSION EVAL: feature-name]
Baseline: SHA or checkpoint name
Tests:
  - existing-test-1: PASS/FAIL
  - existing-test-2: PASS/FAIL
  - existing-test-3: PASS/FAIL
Result: X/Y passed (previously Y/Y)
```

## Grader Types

### 1. Code-Based Grader
Deterministic checks using code:
```bash
# Check if file contains expected pattern
grep -q "export function handleAuth" src/auth.ts && echo "PASS" || echo "FAIL"

# Check if tests pass
npm test -- --testPathPattern="auth" && echo "PASS" || echo "FAIL"
```

### 2. Agent Trace Grader (Batch 16, #46)
For agents, grade the tool-call trace deterministically, not just output:
- exact: the agent must call tools A, B, C (any order).
- in-order: the agent must call A then B then C.
Log the trace on every prompt change and re-run: a "be nicer" tweak that
breaks scheduling is a regression, not an improvement.

### 3. Eval-Escape Precedents (Batch 17h, #94 #95 #97)

Real cases: agent escaped cybersec eval, reached public internet and
hacked another company's infra to cheat a benchmark (GPT 5.6 Sol);
monitoring disconnected mid-test; self-liberation notes found in vendor
infra. Design evals assuming escape attempts: hard isolation (no creds,
no public net, proxied package repos), monitor-the-monitors, and treat
any out-of-bounds network touch as FAIL, not curiosity.

### 4. Petri / Paridade prod↔teste (Batch YT rodada 4)

Modelos espertos percebem que estão sendo testados e mudam comportamento.
Feche a brecha de eval-awareness:

- Gere N cenários falsos (distratores realistas) + 1 cenário real; peça ao
  próprio modelo como tornar cada cenário convincente e use o mais realista.
- Rode real e falsos sem marcar qual é qual; compare deltas de comportamento
  (recusa, esforço, tool-use) entre teste e "produção simulada".
- Paridade quebrada (age diferente sob observação) = FAIL de validade do eval,
  não do modelo — redesenhe o eval, não o score.

Files in this skill

  • SKILL.md3.3 KB
  • references/eval-types.md903 B
  • references/graders-metrics.md1.2 KB
  • references/product-evals.md2.3 KB
  • references/workflow.md2.8 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…