Skip to content
Back to skills

Ai Evals

ASecurity

Designs trustworthy LLM, agent, responsible-AI, and multimodal evaluations. Use when measuring quality, fairness, privacy, grounding, safety, or judge reliability.

  • 89 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 2, 2026
ai-agentspythonrustgorailstestingapici/cdperformance

Works with

  • claude code
  • api

Security analysis

A100/100

Pro scans all 18 files and shows the line behind each finding

Scanned October 6, 2026

npx -y skills add vasilyu1983/AI-Agents-public --skill ai-evals --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Ai Evals?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Ai Evals
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vasilyu1983-ai-evals/badge)](https://www.skillsdirectory.com/skills/vasilyu1983-ai-evals)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: ai-evals
description: "Designs trustworthy LLM, agent, responsible-AI, and multimodal evaluations. Use when measuring quality, fairness, privacy, grounding, safety, or judge reliability."
compatibility: Portable core. Works on Claude Code and Codex.
version: "1.8"
last_validated: 2026-08-24
---

# AI Evaluation and Fine-Tuning Methodology Skill

Use this skill to calibrate graders, derive release gates, and compare evaluation
or adaptation methods. Domain skills own what to measure; this skill owns the
measurement controls. Training loss is telemetry; held-out behavior determines
whether an adaptation improved the system.

## Quick Reference

| Task | Read or Run | Outcome |
|------|-------------|---------|
| Build a (question, ideal-answer) set and tune it | `references/dataset-construction.md` | Sourcing, ideal-answer authoring, run→compare→tune loop, robustness slice by perturbing memorized cases |
| Compose an eval dataset | `references/eval-dataset-design.md` | Task distribution, difficulty, sampling, annotation, versioning, contamination, bias audits, golden vs dynamic sets |
| Stop a judge from rating its own output high | `references/llm-judge-bias.md` | Self-preference, position, length, verbosity controls |
| Pick / wire an eval framework | `references/framework-integration.md` | inspect-ai, lighteval, Ragas, DeepEval, promptfoo, Braintrust integration snippets + when to use each |
| Choose metrics and a pass threshold defensibly | `references/threshold-derivation.md` | Metric set by use case with direction and action-if-fails; derive thresholds from a labeled set; inter-rater agreement; gate design |
| Stop flaky runs reading as regressions | `references/flake-and-reproducibility.md` | pass@k vs pass^k, seeds, temperature, quarantine, contamination/leakage, system-benchmark hazards (hardware lottery, thermal/background load, unreported CIs) |
| Decide if "A beats B" is real, size the set | `references/eval-statistics.md` | Bootstrap CIs, McNemar, power/MDE sizing, FDR, variance reduction |
| Get maximum from an LLM | `references/llm-optimization-technique-map.md` | Technique ladder across prompts, data, RAG/tools, test-time compute, SFT, preference/RFT, PEFT, distillation; measured case where a confidence scorer scored below no scorer |
| Decide whether and how to fine-tune | `references/fine-tuning-eval-loop.md` | Prompt/RAG/tool baseline, SFT vs preference/RFT vs PEFT, split hygiene, promotion gates |
| Evaluate on live/production traffic | `references/online-production-eval.md` | Offline-online correlation, A/B+guardrails, shadow/canary, drift, regression replay, HITL |
| Turn user behavior into eval and preference data | `references/conversational-feedback-signals.md` | Implicit NL signals, action signals, edit→preference pairs, collection timing, feedback biases |
| Evaluate refusals, jailbreaks, harm | `references/safety-redteam-eval.md` | Over/under-refusal, ASR per attack family, injection, harm rubrics, robustness |
| Evaluate responsible or multimodal AI | `references/responsible-multimodal-evaluation.md` | Intersectional fairness, privacy/memorization/poisoning, oversight/appeals, provenance, grounding, diffusion, multimodal attacks, latency/cost |
| Go beyond one judge | `references/advanced-judging.md` | Juries, fine-tuned judges, CoT/probability scoring, calibration (kappa/ECE), agentic reward, human-vs-judge routing |
| Evaluate a retrieval, KB, or memory system | `references/retrieval-and-memory-eval.md` | Independent held-out set, slice floors, Wilson CIs, confident-wrong rate, per-stage tests, change-triggered regression gate, state-maintenance acceptance tests, cross-tenant canaries, learned-memory baselines, vendor benchmark claims |
| Measure whether memory skills improve an agent | [memory-skills-ablation](references/memory-skills-ablation.md) | Validate paired skills-on/off outcomes, complete safety slices, configuration parity and measured telemetry; export descriptive paired results |
| Measure hallucination | `references/hallucination-eval.md` | Taxonomy, claim-level scoring, abstention next to hallucination rate, reference-free triage |

## Scope Boundaries (Use These Skills for Depth)

- **Domain metrics for retrieval** (nDCG/MRR/recall, faithfulness) -> [ai-rag](../ai-rag/SKILL.md)
- **Agent harness: tool/trace grading, refusal/red-team packs, regression suites, CI status model** -> [qa-agent-testing](../qa-agent-testing/SKILL.md)
- **Coding-agent golden tasks, trace replay, cost ops** -> [ai-coding-agents-observability-evals](../ai-coding-agents-observability-evals/SKILL.md)
- **Running benchmark harnesses on Hub models** -> use the `huggingface-skills:` plugin (external)
- **Prompt CI/CD and structured output contracts** -> [ai-prompt-engineering](../ai-prompt-engineering/SKILL.md)
- **Model selection** -> [ai-architecture-advisor](../ai-architecture-advisor/SKILL.md);
  **serving and quantization** -> [ai-llm-inference](../ai-llm-inference/SKILL.md);
  **API cost economics** -> [ops-cost-optimization](../ops-cost-optimization/SKILL.md)

## Workflow

1. **Transform the vague ask into a verifiable goal.** "Is it good?" is not
   gradeable. Ask: which case would fail first if the requirement reverted?
2. **Build the dataset before the grader.** Source real questions, author ideal
   answers from the system's allowed context, and plan the run→compare→tune loop
   — see `references/dataset-construction.md`. No dataset, no eval.
3. **Pick the cheapest grader that works.** Deterministic check > LLM judge >
   human. Reserve the LLM judge for what code cannot decide (Rule 5: use the
   model only for judgment calls).
4. **If using an LLM judge, control its bias** before trusting any number — see
   `references/llm-judge-bias.md`. Judge bias can confound quality with surface features. When one judge isn't enough (high stakes, weak
   agreement, open-ended), escalate to juries / fine-tuned judges / calibrated
   scoring — see `references/advanced-judging.md`.
5. **Control flake and leakage** with repeated trials (pass^k for reliability
   gates, pass@k only for capability), low judge temperature, seed
   pinning, and held-out testsets — see `references/flake-and-reproducibility.md`.
6. **Derive thresholds from a labeled calibration set**, not intuition or copied
   targets — see `references/threshold-derivation.md`. **Size the gating set and
   judge "A beats B" with statistics** (bootstrap CIs, McNemar, power/MDE, FDR) —
  see `references/eval-statistics.md`. For frozen paired results, run
  `python3 scripts/analyze_paired_results.py results.csv --estimand unit_mean`;
  declare whether units or clusters define the target population first. A population-level superiority claim needs design-compatible uncertainty;
  exact fixed-suite differences remain descriptive evidence.
7. **Only fine-tune after the baseline has earned it.** Compare prompt/RAG/tool
   fixes first, then choose SFT for imitation/style/format/tool-call behavior,
   preference/RFT for rubric-scored reasoning or tradeoffs, and PEFT/LoRA/QLoRA
   when adapting an open model under compute or deployment constraints — see
   `references/fine-tuning-eval-loop.md`. Training loss is telemetry; held-out
   behavior is the verdict.
8. **Apply the full optimization ladder, not one pet method.** For maximum LLM
   performance, evaluate cheap prompt/context/tool fixes, then inference-time
   methods (self-consistency, best-of-N, rerank/verify/refine), then data/SFT,
   preference/RFT/RLVR, PEFT, and distillation as the evidence warrants — see
   `references/llm-optimization-technique-map.md`. Each technique gets its own
   failure mode and gate.
9. **Gate loudly.** A gate that passes while silently skipping cases is a
   failure dressed as success (Rule 12: fail loud). Report skipped/quarantined
   cases in the gate output.
10. **Extend past the offline gate where the system warrants it.** Add
   safety/red-team evaluation (refusal precision/recall, jailbreak ASR,
   injection, harm rubrics) — see `references/safety-redteam-eval.md` — and, once
   in production, online evaluation (offline-online correlation, A/B with
   guardrails, drift, regression replay) — see `references/online-production-eval.md`.
   The offline gate is a filter; production is the verdict.

## Core Principles

- **The judge is a model with failure modes.** Treat its scores as one
  calibrated input, never as ground truth.
- **Judge independence needs evidence.** Prefer a different judge model and
  calibrate against human labels and planted failures. A different family reduces
  one self-preference risk; it does not establish an unbiased grader.
- **Behavior, not plausibility.** Rubrics that reward "looks good" reward length
  and confidence. Pin rubrics to verifiable behavior.
- **No threshold without a labeled set.** A copied target (">95%") is a guess
  until validated on your own distribution.
- **No fine-tune without a baseline and a holdout.** A tuned model that beats no
  prompt/RAG/tool baseline, or only wins on the training/dev set, has not earned
  release.
- **No "maximum performance" without a technique ladder.** The best result often
  comes from composition: cleaner data + stronger retrieval/tool contracts +
  calibrated judge + small test-time search + selective post-training. Test the
  cheapest credible lift before moving weights.
- **Optimize behavior, not hidden knowledge.** Fine-tune for stable formatting,
  domain style, tool-use patterns, rubric-following, or compact specialized
  behavior. Use retrieval/context for facts that change or must be cited.
- **Locate variability before changing the gate.** Regrade frozen outputs to
  separate judge instability from system unreliability; keep real stochastic
  failures in the reliability denominator.
- **Held-out or it's contaminated.** If tuning ever saw the eval cases, the
  scores are inflated.
- **Check that the proxy still tracks the outcome.** Tuning to a metric can
  weaken its relationship with user outcomes. Re-anchor it against fresh human
  labels or production evidence when the model, grader, or traffic changes.
- **Match uncertainty to the claim.** Population-level pass rates, win rates,
  and judge-human agreement need design-compatible uncertainty. Exact counts
  over a fixed deterministic suite describe that suite and need their
  denominator and limits; do not invent sampling intervals for them. See
  `references/eval-statistics.md`.

- Eval framework APIs (inspect-ai, lighteval, Ragas, DeepEval, promptfoo,
  Braintrust) change across releases. Verify current API and version against
  official docs before recommending a specific call or flag.
- Optimization-method papers from arXiv are often preprint-only and
  benchmark-sensitive. Treat unreplicated methods as `validate`, not `promote`,
  until they beat a strong local baseline with cost/latency/safety gates.
- Judge-bias findings (position, length, self-preference) are well-replicated,
  but their magnitudes (and for length, the direction) are judge- and
  prompt-dependent — re-measure on your own setup; do not quote a fixed number as universal.
- Fine-tuned-judge models (Prometheus, JudgeLM, and successors) and jailbreak
  attack/defense results move fast — verify the current model, license, and
  benchmark-agreement claims before recommending a specific judge or asserting a
  model is robust to a given attack family.
- Indirect prompt injection is a leading agentic attack class; treat any
  "the agent is safe against injection" claim as requiring fresh adaptive testing.
- Statistics methods (bootstrap, McNemar, FDR, power/MDE) are stable, but verify
  the exact API (scipy/statsmodels) before copying a call.

## Release Decision Gate

### Select foundations only when they change the eval decision

- If a proxy score is being interpreted as quality, trust, fairness, or another
  construct, or compared across populations or grader versions, use
  [measurement theory](../foundations-measurement-theory/SKILL.md). Return the
  construct, observed proxy, interpretation evidence, comparability limits,
  and missing validation. Skip this audit for a routine execution of an already
  documented instrument with unchanged interpretation and population.
- If uncertainty, dependent repeats/clusters, multiple comparisons, or repeated
  stopping can change a conclusion, use
  [statistical inference](../foundations-statistical-inference/SKILL.md). Return
  the estimand, independent unit, interval method and assumptions, and the
  multiplicity/stopping rule. Skip inference for reporting exact outcomes of a
  fixed deterministic fixture suite without a population claim; report its
  denominator and scope instead. Neither foundation substitutes for this skill's
  dataset, grader controls, or release decision.

Write the release rule before scoring: required tasks and slices, minimum acceptable effect, uncertainty method, regression budget, and handling for grader errors or missing results. A candidate passes only when every blocking slice is present and the predeclared decision rule clears. Learned, model, heuristic, and human graders must pass planted-good and planted-bad controls that detect calibration drift; deterministic exact-match, schema, executable-test, and invariant graders instead need versioned implementations with relevant positive and negative fixtures. Missing denominators, stale fingerprints, failed grader controls, or an unavailable judge produce `inconclusive`, never an implicit pass. For memory-skill effectiveness, use the [paired ablation procedure](references/memory-skills-ablation.md); public development fixtures validate the harness, while an independently frozen set is required for a held-out claim. Keep measured results separate from the release recommendation.

## Known Traps

- Grading an agent with the same model that produced the output (self-preference)
- Comparing two candidates in fixed order and trusting the winner (position bias)
- Copying a `>95%` threshold from a blog without validating it on your data
- One judge call per request with no cheap deterministic pre-filter (cost blowup)
- Quarantining genuine stochastic system failures until the gate becomes green
- Generating a synthetic testset from the same docs used to tune the system
- Reporting a builder-written eval as held-out for a retrieval/KB/memory system, or
  letting the eval fall back silently when the index is stale (see
  `references/retrieval-and-memory-eval.md`)
- Reporting "all passed" when some cases were skipped or errored (silent success)
- Claiming "A beats B" from a point estimate with no confidence interval or test
- Gating a small regression on a set far too small to detect it (no power check)
- Tuning safety to block harm without a benign set, so the model over-refuses
- Trusting a seed for reproducibility through a hosted API that isn't deterministic
- Fine-tuning because the prompt is messy, the retrieval is broken, or the tool
  contract is ambiguous
- Declaring the fine-tune better from training loss, validation loss, or one
  cherry-picked demo instead of a paired held-out eval with CIs
- Letting the training set, grader calibration set, and release gate share cases
- Training a judge or reward model on labels produced only by the same model
  family it will later grade

## Navigation

Resources:

- [references/dataset-construction.md](references/dataset-construction.md) - Sourcing questions, authoring ideal answers, run→compare→tune loop
- [references/eval-dataset-design.md](references/eval-dataset-design.md) - Dataset composition: distribution, difficulty, sampling, annotation, versioning, contamination, bias, golden vs dynamic
- [references/llm-judge-bias.md](references/llm-judge-bias.md) - Judge bias taxonomy and controls
- [references/framework-integration.md](references/framework-integration.md) - Framework selection and integration snippets
- [references/threshold-derivation.md](references/threshold-derivation.md) - Metric set by use case, deriving thresholds and gates from labeled data
- [references/flake-and-reproducibility.md](references/flake-and-reproducibility.md) - Flake, seeds, contamination, leakage, system-benchmark hazards
- [references/eval-statistics.md](references/eval-statistics.md) - Paired cluster bootstrap workflow and analyzer, McNemar, power/MDE, FDR, variance reduction
- [references/llm-optimization-technique-map.md](references/llm-optimization-technique-map.md) - Maximum-performance technique ladder and eval gates
- [references/fine-tuning-eval-loop.md](references/fine-tuning-eval-loop.md) - Eval-first fine-tuning decisions, SFT/preference/RFT/PEFT selection, split hygiene, promotion gates
- [references/online-production-eval.md](references/online-production-eval.md) - Offline-online correlation, A/B+guardrails, drift, replay, HITL
- [references/conversational-feedback-signals.md](references/conversational-feedback-signals.md) - Implicit/action user-feedback taxonomy, edit→preference pairs, collection timing, feedback biases
- [references/safety-redteam-eval.md](references/safety-redteam-eval.md) - Refusal precision/recall, jailbreak/injection, harm rubrics, robustness
- [references/advanced-judging.md](references/advanced-judging.md) - Juries, fine-tuned judges, scoring methods, calibration, human-vs-judge routing, agentic reward
- [references/retrieval-and-memory-eval.md](references/retrieval-and-memory-eval.md) - Independent held-out sets, slice floors, per-stage pipeline tests, change-triggered regression gate, exact-state scoring
- [references/hallucination-eval.md](references/hallucination-eval.md) - Hallucination taxonomy, claim-level metrics, abstention, judge prompt, reference-free triage
- [references/responsible-multimodal-evaluation.md](references/responsible-multimodal-evaluation.md) - Responsible-AI measurement and multimodal grounding, diffusion, safety/red-team, latency, and cost gates
- [scripts/prompt_eval_runner.py](scripts/prompt_eval_runner.py) - Offline regression check of pre-collected outputs in a JSONL suite against `expected_substrings` and/or `expected_schema`. A record with no assertions fails; malformed assertion fields, non-JSON numeric constants, and duplicate case IDs exit 2. It calls no model API and measures nothing statistical; compare two variants with `scripts/analyze_paired_results.py`.
- [data/sources.json](data/sources.json) - Sources to verify against

Related skills:

- [ai-llm](../ai-llm/SKILL.md), [ai-rag](../ai-rag/SKILL.md), [ai-coding-agents-observability-evals](../ai-coding-agents-observability-evals/SKILL.md), [qa-agent-testing](../qa-agent-testing/SKILL.md)

## Learnings Loop

Consult `learnings.consolidated.md` for relevant prior decisions or pitfalls;
open `learnings.md` only when their history is needed. Skip both when unrelated
to the task. After applying it, append one dated bullet to
`learnings.md` via `agents-skills-feedback-loop/scripts/append_learning.py` if you
hit a pattern, mistake, or surprising fact. Do not modify `SKILL.md` itself.

Files in this skill

  • SKILL.md17.7 KB
  • agents/openai.yaml348 B
  • data/sources.json39.3 KB
  • learnings.consolidated.md583 B
  • learnings.md2.5 KB
  • references/advanced-judging.md9.3 KB
  • references/conversational-feedback-signals.md11.9 KB
  • references/dataset-construction.md18.1 KB
  • references/eval-statistics.md6.2 KB
  • references/fine-tuning-eval-loop.md9 KB
  • references/flake-and-reproducibility.md7.1 KB
  • references/framework-integration.md5.8 KB
  • references/llm-judge-bias.md5.9 KB
  • references/llm-optimization-technique-map.md19.2 KB
  • references/online-production-eval.md6.5 KB
  • references/responsible-multimodal-evaluation.md13.5 KB
  • references/safety-redteam-eval.md11.4 KB
  • references/threshold-derivation.md4 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…