Skip to content
Back to skills

Agent Evaluation

ASecurity

Evaluate LLM agents and tool-using workflows—task success, tool accuracy, latency/cost, safety, and regression suites. Use when shipping agent features, comparing prompts/models, or debugging agent failures. Triggers: "agent eval", "benchmark agent", "tool accuracy", "agent regression".

  • 25 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added May 27, 2026
data-aigodebugging

Security analysis

A100/100

Scanned May 27, 2026

npx -y skills add charlieviettq/awesome-agent-skill --skill agent-evaluation --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Agent Evaluation?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Agent Evaluation
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/charlieviettq-agent-evaluation/badge)](https://www.skillsdirectory.com/skills/charlieviettq-agent-evaluation)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: agent-evaluation
description: >
  Evaluate LLM agents and tool-using workflows—task success, tool accuracy,
  latency/cost, safety, and regression suites. Use when shipping agent features,
  comparing prompts/models, or debugging agent failures.
  Triggers: "agent eval", "benchmark agent", "tool accuracy", "agent regression".
---

# Agent evaluation

## What to measure

| Dimension | Examples |
|-----------|----------|
| Task success | End state matches spec (binary or rubric) |
| Tool use | Correct tool, valid args, no spurious calls |
| Safety | No policy violations, no secret leakage |
| Efficiency | Tokens, latency, tool call count |
| Stability | Same input -> consistent outcome across runs |

## Workflow

1. **Define tasks** — realistic user intents with clear pass/fail or scored rubric.
2. **Build dataset** — golden set + edge cases (errors, ambiguous input, empty context).
3. **Run baseline** — fixed model/settings; log traces (inputs, tools, outputs).
4. **Score** — automated checks first; human review for ambiguous cases.
5. **Compare** — A/B prompts, models, or tool schemas; report deltas with confidence notes.
6. **Gate** — block release on regression in must-pass tasks.

## Automated checks

- Schema validation on tool arguments.
- Assert final answer contains required fields or avoids forbidden content.
- Snapshot tests for deterministic sub-steps where possible.

## Human rubric (when needed)

Score 1-5 on: correctness, completeness, tone, safety. Document disagreements.

## Anti-patterns

- Eval only on cherry-picked happy paths.
- Changing task and model simultaneously without isolation.
- No trace logs when debugging tool failures.

## Output

Summary table: variant | success rate | avg tools | avg latency | notes.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…