Back to skills
SKILL.md
Llm Evaluation
ASecurityLLM evaluation — automated metrics, human feedback, benchmarking. Use when testing performance, measuring AI quality, or establishing evaluation frameworks.
- 3 stars
- 0 votes
- 0 copies
- 3 views
- Added May 28, 2026
Security analysis
100/100npx -y skills add martineserios/thebrana --skill llm-evaluation --agent claude-codeAre you the author of Llm Evaluation?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/martineserios-llm-evaluation)---
name: llm-evaluation
description: "LLM evaluation — automated metrics, human feedback, benchmarking. Use when testing performance, measuring AI quality, or establishing evaluation frameworks."
group: brana
keywords: [llm, evaluation, eval, testing, llm-judge, a-b-testing, benchmarking, bleu, rouge, bertscore, langsmith]
allowed-tools: [Read, Glob, Grep, AskUserQuestion]
status: experimental
source: "https://github.com/wshobson/agents @llm-evaluation"
acquired: "2026-04-30"
quarantine: true
---
<!-- PROCEDURE_FILE: procedures/llm-evaluation.md -->
Read and execute the full procedure from `system/procedures/llm-evaluation.md`.
> QUARANTINE: Community-tier skill. Read-only tools only. Verify patterns against official docs before applying.
Attribution
Comments
Loading comments…