Skip to content
Back to skills

Harness Eval

BSecurity

Structured SE task evaluation using 15 benchmark definitions from claude-code-harness research

  • 43 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 9, 2026
businessgoshellbashgitapidatabasesecurityperformance

Works with

  • cli
  • api

Security analysis

B85/100
  • highPerforms destructive filesystem operations

Pro shows the line behind each finding and how to fix it

Scanned September 9, 2026

npx -y skills add baekenough/oh-my-customcode --skill harness-eval --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Harness Eval?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Harness Eval
[![Security: B — Skills Directory](https://www.skillsdirectory.com/api/skills/baekenough-harness-eval/badge)](https://www.skillsdirectory.com/skills/baekenough-harness-eval)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: harness-eval
description: Structured SE task evaluation using 15 benchmark definitions from claude-code-harness research
scope: harness
user-invocable: true
argument-hint: "[--preset all|quick] [--task task-name]"
effort: high
version: 1.0.0
---

# Harness Eval — Structured SE Task Benchmark

## Purpose

Evaluate agent quality using 15 structured software engineering task definitions with quantitative scoring. Based on research from [revfactory/claude-code-harness](https://github.com/revfactory/claude-code-harness) which demonstrated 60% improvement (49.5 → 79.3 points) through structured pre-configuration.

## Usage

```
/harness-eval                    # Run all 15 benchmarks
/harness-eval --preset quick     # Run top 5 high-impact benchmarks
/harness-eval --task api-design  # Run specific task benchmark
```

## Quality Dimensions

| Dimension | Weight | Description |
|-----------|--------|-------------|
| Test Coverage | 30% | Unit test count, edge case coverage, assertion quality |
| Architecture Design | 25% | Separation of concerns, dependency management, scalability |
| Error Handling | 25% | Input validation, error propagation, recovery strategies |
| Extensibility | 20% | Plugin points, configuration flexibility, API surface |

## 15 SE Task Benchmark Suite

| # | Task | Category | Key Evaluation Criteria |
|---|------|----------|------------------------|
| 1 | API Design | Architecture | RESTful conventions, versioning, error responses |
| 2 | Data Modeling | Architecture | Schema normalization, relationships, indexing |
| 3 | Authentication Flow | Security | Token management, session handling, OWASP compliance |
| 4 | Test Suite Creation | Quality | Coverage breadth, assertion quality, edge cases |
| 5 | Error Handler | Reliability | Error classification, recovery, user feedback |
| 6 | Logging System | Observability | Structured logging, levels, correlation IDs |
| 7 | Configuration Manager | Operations | Env-based config, validation, secrets handling |
| 8 | CLI Tool | UX | Argument parsing, help text, exit codes |
| 9 | Database Migration | Data | Reversibility, data preservation, zero-downtime |
| 10 | Cache Layer | Performance | Invalidation strategy, TTL, cache-aside pattern |
| 11 | Queue Consumer | Reliability | Idempotency, retry logic, dead letter handling |
| 12 | Middleware Chain | Architecture | Composability, ordering, short-circuiting |
| 13 | File Processor | I/O | Streaming, error recovery, format validation |
| 14 | Webhook Handler | Integration | Signature verification, retry tolerance, idempotency |
| 15 | Rate Limiter | Security | Algorithm choice, distributed state, fairness |

## Scoring Rubric

Each task is scored 0-100 across the 4 quality dimensions:

```
Score = (test_coverage × 0.30) + (architecture × 0.25) + (error_handling × 0.25) + (extensibility × 0.20)
```

### Score Thresholds

| Score Range | Grade | Interpretation |
|-------------|-------|----------------|
| 80-100 | A | Production-ready, well-structured |
| 60-79 | B | Functional with minor gaps |
| 40-59 | C | Works but needs improvement |
| 0-39 | D | Significant structural issues |

## Presets

### `all` (default)
Run all 15 tasks. Full evaluation ~45 minutes.

### `quick`
Run top 5 high-impact tasks (1, 3, 4, 5, 12). Quick evaluation ~15 minutes.

## Integration with evaluator-optimizer

This skill provides preset rubrics for the evaluator-optimizer pipeline:

```
/harness-eval → loads rubric → evaluator-optimizer executes → scoring → report
```

The evaluator-optimizer skill's `pre_negotiation` phase accepts harness-eval rubric dimensions as sprint contract criteria.

## Output

Results saved to `.claude/outputs/sessions/{YYYY-MM-DD}/harness-eval-{HHmmss}.md` with per-task scores and aggregate grade.

### Tool: Writing artifacts under .claude/outputs/

Under `mode: "bypassPermissions"`, direct Write/Edit/Bash on `.claude/**` paths (including `.claude/outputs/sessions/`) is permitted (CC v2.1.121+, #1101) — no `/tmp` wrapping is needed.

Write harness-eval results directly to `.claude/outputs/sessions/$(date +%Y-%m-%d)/harness-eval-$(date +%H%M%S).md`. Read-only Bash on `.claude/outputs/` (e.g., `cat`, `head`, `wc`) is allowed for verification. Catastrophic shell operations (e.g., `rm -rf /`) remain blocked by independent safety guards. For CC < v2.1.121, see git history for the legacy `/tmp/*.sh` bypass pattern.

Reference: R006 Sensitive Path Handling, R010 Universal bypassPermissions, #1101.


## 4-Metric Quantitative Layer (added v0.113.0, #1025)

The 15 benchmark tasks defined here measure **task correctness** (pass/fail). For agent efficiency comparison and trajectory analysis, layer the 4-metric framework on top:

- **correctness** — existing benchmark pass/fail (unchanged)
- **step_ratio** — observed_steps / ideal_steps (efficiency)
- **tool_call_ratio** — observed_tool_calls / ideal_tool_calls (efficiency)
- **latency_ratio** — observed_latency / ideal_latency (efficiency)

### Workflow
1. Run benchmark task → collect correctness result + trajectory (steps, tool_calls, latency)
2. Compare against ideal trajectory annotation (see `agent-eval-framework` skill)
3. Phase 1 gate: correctness MUST pass; Phase 2 gate: efficiency comparison among passing variants

### Ideal Trajectory per Benchmark
For each of the 15 benchmark tasks, an ideal trajectory should be authored. Annotation schema:
```yaml
task_id: <benchmark-id>
capability: <category>
ideal:
  steps: <int>
  tool_calls: <int>
  latency_seconds: <float>
```

### Cross-references
- Skill: `agent-eval-framework` (4-metric framework definition)
- Guide: `guides/agent-eval/README.md` (measurement methodology)
- Issue: #1025

## Attribution

Evaluation framework based on research by [revfactory/claude-code-harness](https://github.com/revfactory/claude-code-harness). Adapted for oh-my-customcode's evaluator-optimizer pipeline with permission.

## Related Guide

- `guides/harness-engineering/` — 하네스 엔지니어링 통합 가이드 (Benchmark Evaluation Layer 관점에서 harness-eval 위치)

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…