Use when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual recall) or agent trajectories (tool correctness, completion), or picking an eval framework. NOT building the agent loop, tools or RAG plumbing (that is `building-agents`).
Installs into .claude/skills of the current project.
Are you the author of Agent Eval?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/ericrisco-agent-eval)
---
name: agent-eval
description: "Use when measuring whether an LLM or agent system actually got better and gating merges on it: golden sets, fixing an inflated LLM-as-judge, scoring RAG (faithfulness, contextual recall) or agent trajectories (tool correctness, completion), or picking an eval framework. NOT building the agent loop, tools or RAG plumbing (that is `building-agents`)."
tags: [evals, llm, agents, llm-as-judge, regression-gate, ai]
recommends: [building-agents, prompt-engineering, observability]
origin: risco
---
# Measure agent quality you can defend and gate on
Turn "the agent feels better" into a number you can put in a PR check. You own the eval dataset, the scorer mix, the LLM-as-judge calibration, and the block-on-regression CI gate — framework-neutral, provider-neutral.
## Do NOT use — route instead
| The ask | Route to | Why it is not this skill |
| --- | --- | --- |
| Build the agent loop, tools, RAG plumbing | `building-agents` | It builds the system; you score it. They cross-link. |
| "Make the answers shorter / rewrite the prompt" | `prompt-engineering` | Evals say it is worse; that skill changes the words. You never edit the prompt. |
| pytest/jest on deterministic functions | `testing-py` / `testing-web` | Assert-equals on pure code, not stochastic outputs scored by a judge. |
| Dashboards / tracing of live production traffic | `observability` | Online monitoring; you are offline + pre-merge. |
| Red-team, jailbreak, prompt injection | `agent-safety` | Adversarial coverage, not quality measurement. |
| Per-token cost budgets and accounting | `cost-tracking` | You report cost-per-task as one metric; the discipline lives there. |
| A/B stats on product/funnel metrics | `ab-testing` | Web experiments, not offline model comparison on a fixed set. |
## The eval anatomy
Every framework instantiates the same five-stage pipeline. Learn it once; the tool is a detail.
```text
dataset ──▶ runner ──▶ scorers ──▶ metrics ──▶ gate
(JSONL (calls the (det / judge (aggregate + (pass/fail
golden system per / human) bootstrap CI) exit code)
set) case)
```
DeepEval, Inspect AI, and promptfoo are all just opinionated wrappers around this. If you understand the stages you can switch tools without relearning the craft.
## Build the dataset first
Build your own golden set — a public leaderboard number is not your number, because identical model weights swing SWE-Bench Verified by 10–20 points just by changing the harness. Measure *your* task on *your* data. The dataset is the asset; everything else is replaceable. Rules:
- **50–200 hand-labeled cases per failure mode**, not per total. Coverage of how the system fails beats raw volume. 80 real failure cases > 1000 generic ones.
- **Never synthetic-only.** A set the model wrote will not surface the model's blind spots. Mine real traffic / tickets / transcripts and hand-label.
- **Version it in git as JSONL**, a first-class reviewed asset — same as code. Diffs are reviewable; relabels are auditable.
- **Decontaminate.** The eval set must not appear in training data or few-shot examples, or the score is a memorization artifact, not a capability.
Case schema — one JSON object per line:
```jsonl
{"id":"refund-001","input":"Where is my refund for order 4821?","expected":"States refunds take 5-7 business days and asks for nothing already on file","context":["policy: refunds 5-7 business days"],"meta":{"failure_mode":"hallucinated_policy","source":"ticket#4821"}}
{"id":"refund-002","input":"Cancel my subscription and refund this month","expected":"Cancels, refunds prorated amount, confirms no future charge","context":["policy: prorated refund on cancel"],"meta":{"failure_mode":"missed_tool_call","source":"ticket#5190"}}
```
`failure_mode` in `meta` is what lets you slice metrics by mode and find *which* kind of bug regressed — not just that the aggregate dropped.
> Bad: "generate 1000 test questions with GPT and use those." Good: "80 real failure-mode cases pulled from support tickets, hand-labeled, tagged by failure mode."
## Choose the scorer — the 60/30/10 mix
Reach for the cheapest scorer that correlates with human judgment. Default mix:
| Share | Scorer kind | Use for | Why |
| --- | --- | --- | --- |
| ~60% | Deterministic — exact match, regex, JSON-schema validation, latency threshold | Anything with a checkable shape: format, required fields, a known string, a budget | Free, instant, zero drift. Never spend a judge call on something a regex settles. |
| ~30% | LLM-as-judge — G-Eval, DAG, custom Python scorer | Meaning: is this answer faithful, relevant, helpful | Only where correctness is semantic. Costs money and can drift — so calibrate it. |
| ~10% | Human-in-the-loop | Genuinely ambiguous cases the judge disagrees on | The ground truth you calibrate the judge against. |
One `Scorer` protocol, two implementations behind it — deterministic and judge are interchangeable to the runner:
```python
from typing import Protocol
class Scorer(Protocol):
name: str
def score(self, case: dict, output: str) -> float: ... # 0.0–1.0
class JsonSchemaScorer:
name = "schema_valid"
def score(self, case, output): # deterministic, free, no drift
import json
try:
json.loads(output)
return 1.0
except ValueError:
return 0.0
class FaithfulnessJudge:
name = "faithfulness"
def __init__(self, judge_model): self.judge = judge_model
def score(self, case, output): # judge only where meaning matters
return self.judge.rate(case["context"], output) # see judge-design.md
```
## LLM-as-judge you can trust
A score you do not trust is worse than no score: an uncalibrated judge gives false confidence, which is more dangerous than admitted ignorance. Each rule, with its why:
- **Judge model ≥ system under test.** A weaker judge cannot reliably rank a stronger system — it scores noise.
- **The rubric must force a written rationale before the score.** Rationale-first judging is what pushes judge–human agreement to ~85% — higher than two humans agree with each other. A bare number is a vibe with a decimal point.
- **Pairwise beats pointwise for stability.** "Is A or B better?" is more reproducible than "rate A from 1–10," which inflates and clusters at 8–9.
- **Swap positions and average.** Judges favor whichever answer came first; run A-then-B and B-then-A to cancel position bias.
- **Calibrate against human gold and report the agreement before you gate anything on the judge.** Not a formality — this is the step that makes every number downstream defensible.
> Bad judge prompt: "Rate this answer 1–10." → everything lands 8–9, useless.
> Good: "Compare answer A and answer B against the reference. First write one sentence on each per the rubric, then output the better label." → forces reasoning, gives a stable signal.
Full rubric templates (pointwise + pairwise), the position-swap harness, the calibration script (agreement / Cohen's kappa vs human gold), G-Eval vs DAG, and the judge bias catalog (length, position, self-preference) with mitigations live in **[references/judge-design.md](references/judge-design.md)**.
## Agent and RAG scorers
Score the path, not only the destination. Beyond exact/judge:
**RAG** (DeepEval / RAGAS names):
- **Faithfulness** — does the answer only claim what the retrieved context supports? Catches hallucination.
- **Answer relevancy** — does it actually address the question, or drift?
- **Contextual recall / precision** — did retrieval fetch the right chunks, and not bury them in noise? Separates a retrieval bug from a generation bug.
**Agent:**
- **Tool correctness** — right tool, right arguments, right order.
- **Task completion / goal accuracy** — did it finish the job, not just produce plausible text.
- **Trajectory scoring** — grade the sequence of steps. A correct final answer from a wrong path will fail differently next time; only trajectory scoring catches it.
The system side of these (how the loop and tools are built) is `../building-agents/SKILL.md`; a common system-under-test is `../chatbot/SKILL.md`.
## The regression gate
> Gate policy: **block on regression vs a committed baseline, not on an absolute threshold.** An absolute threshold flaps CI on judge noise and gives no signal on drift; "did this PR make a tracked metric worse than `main`?" is the question that matters.
- Compute a **bootstrap confidence interval** on each metric so judge noise alone does not fail the build — only a drop beyond the CI counts.
- The runner writes `eval-report.json` (metrics, per-failure-mode slices, baseline, pass/fail) and **exits non-zero** on a real regression so the merge is blocked.
```python
import json, sys
def gate(current: dict, baseline: dict, margin: float = 0.0) -> int:
regressed = []
for metric, score in current.items():
if metric in baseline and score < baseline[metric] - margin:
regressed.append((metric, baseline[metric], score))
report = {"metrics": current, "baseline": baseline, "regressed": regressed,
"passed": not regressed}
with open("eval-report.json", "w") as f:
json.dump(report, f, indent=2)
if regressed:
for m, b, c in regressed:
print(f"REGRESSION {m}: {b:.3f} -> {c:.3f}", file=sys.stderr)
return 1
return 0
sys.exit(gate(run_eval(), json.load(open("eval-baseline.json"))))
```
The complete provider-neutral runner (JSONL loader, scorer registry, bootstrap-CI metrics), the GitHub Actions workflow, and side-by-side DeepEval-pytest + Inspect-AI Task/Solver/Scorer versions of the same eval live in **[references/runner-and-gate.md](references/runner-and-gate.md)**.
## Framework cheat-sheet
Pick by where the eval runs and what it must do. Versions as of 2026-06 — re-verify, they rot.
| Tool | What it is | Reach for it when |
| --- | --- | --- |
| **DeepEval** v4.0.3 | pytest-native, 50+ metrics, Decision-Graph (DAG) logic | Your CI is Python/pytest and you want metrics that read like tests. |
| **Inspect AI** v0.3.225 (UK AISI) | dataset→Task→Solver→Scorer, bootstrap CIs, first-class tool-use & trajectory logging, 200+ pre-built evals | Multi-provider, safety-adjacent, or you need real trajectory scoring. |
| **promptfoo** (acquired by OpenAI 2026-03) | CLI + YAML, strong pre-deploy + red-team across 50+ vuln types | Config-driven pre-deploy checks; route the red-team half to `agent-safety`. |
| **Braintrust / LangSmith / Phoenix** v16.0.0 | platforms: annotation, regression tracking, dashboards | You need human annotation queues and historical regression tracking. |
> The two-tool pattern is normal, not over-engineering: a light CI gate (DeepEval / RAGAS / promptfoo) **plus** a platform (Braintrust / LangSmith / Arize) for annotation and history. They share data; different jobs.
## Anti-patterns
| Anti-pattern | Why it bites | Do instead |
| --- | --- | --- |
| Vibes-gating ("feels better, merge it") | No artifact to defend or reproduce | Gate on a number from a committed dataset |
| Synthetic-only dataset | Model-written cases miss the model's blind spots | Hand-label real traffic by failure mode |
| Uncalibrated judge | Confident wrong scores; worse than none | Report agreement vs human gold first |
| Judge weaker than system | Cannot rank a stronger system; scores noise | Judge model ≥ system under test |
| Absolute-threshold gate | Flaps CI on judge noise, blind to drift | Block on regression vs baseline + bootstrap CI |
| Shipping on a leaderboard number | Harness effect = 10–20pt swing | Build your own golden set |
| Scoring only the final answer | A right answer from a wrong path regresses later | Score the trajectory too |
| Never relabeling drifted gold | Stale "truth" silently rots the gate | Review and relabel the golden set on a schedule |
## Project grounding
If the workspace has a `02-DOCS/` harness, record the eval policy in `02-DOCS/wiki/stack/evals.md`: dataset location, scorer mix, gate baseline file, judge model, and the failure modes covered. Follow the harness [`wiki-article-template.md`](../harness/references/wiki-article-template.md) (`type: stack`) and index it in `02-DOCS/wiki/index.md`. This is **recorded, not gated** — skip silently if there is no harness.
## verify.sh
`scripts/verify.sh` is read-only and tool-detecting. It validates that every `*.jsonl` golden set in the project parses and that each line carries the required `id`, `input`, `expected` keys; checks the shape of any `eval-report.json`; and runs `ruff` / `mypy` on example Python and `markdownlint` on docs when those tools are installed. Every missing tool prints a yellow WARN and is skipped — never a failure. An empty or clean target exits 0.