Skip to content
Back to skills

Eval Harness

ASecurity

Evaluation harness for testing agent and skill quality through structured benchmarks, regression tests, and quality scoring.

  • 1,760 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added May 29, 2026
ai-agentsbashtestingperformance

Works with

  • claude code

Security analysis

A100/100

Pro scans all 2 files and shows the line behind each finding

Scanned May 29, 2026

npx -y skills add a5c-ai/babysitter --skill eval-harness --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Eval Harness?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Eval Harness
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/a5c-ai-eval-harness/badge)](https://www.skillsdirectory.com/skills/a5c-ai-eval-harness)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: eval-harness
description: Evaluation harness for testing agent and skill quality through structured benchmarks, regression tests, and quality scoring.
allowed-tools: Read, Write, Edit, Bash, Grep, Glob
---

# Eval Harness

## Overview

Evaluation harness methodology adapted from the Everything Claude Code project. Provides structured frameworks for benchmarking agent performance, testing skill quality, and running regression suites.

## Evaluation Types

### 1. Agent Performance Benchmark
- Define test cases with known-correct outputs
- Run agent against each test case
- Score: accuracy, completeness, relevance
- Compare against baseline performance
- Track performance over time

### 2. Skill Quality Testing
- Verify skill instructions produce expected outcomes
- Test edge cases and boundary conditions
- Measure consistency across multiple runs
- Check for harmful or incorrect outputs
- Validate against ground truth

### 3. Regression Suite
- Collection of previously-passing test cases
- Run after any agent/skill modification
- Flag regressions with before/after comparison
- Maintain pass rate threshold (>= 95%)

### 4. Process Verification
- End-to-end process execution with known inputs
- Verify each phase produces expected outputs
- Check task ordering and dependency satisfaction
- Measure total execution time

## Quality Scoring

### Accuracy Score (0-100)
- Correctness of output vs expected
- Partial credit for partially correct outputs
- Penalty for hallucinated or fabricated content

### Completeness Score (0-100)
- Coverage of required output elements
- Missing sections flagged and scored
- Bonus for useful additional context

### Consistency Score (0-100)
- Run same input 3 times
- Compare outputs for semantic similarity
- Flag inconsistencies

### Composite Score
- (accuracy * 0.4 + completeness * 0.3 + consistency * 0.3)
- Threshold: 80 to pass

## When to Use

- After creating new agents or skills
- After modifying existing agents or skills
- Periodic quality audits
- Before promoting skills to production

## Agents Used

- Used by process-level evaluation orchestrators
- No specific agent dependency (evaluates other agents)

Files in this skill

  • README.md230 B
  • SKILL.md2.1 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…