Skip to content
Back to skills

Llm Judge

ASecurity

AI quality judge that scores agent responses 0-10 across helpfulness, accuracy, completeness, and clarity. Use when evaluating multi-agent output or implementing LLM-as-judge quality gates.

  • 3,819 stars
  • 0 votes
  • 0 copies
  • 4 views
  • Added August 31, 2026
ai-agentsgorails

Security analysis

A100/100

Scanned August 31, 2026

npx -y skills add Atmosphere/atmosphere --skill llm-judge --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Llm Judge?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Llm Judge
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/atmosphere-llm-judge/badge)](https://www.skillsdirectory.com/skills/atmosphere-llm-judge)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: llm-judge
description: "AI quality judge that scores agent responses 0-10 across helpfulness, accuracy, completeness, and clarity. Use when evaluating multi-agent output or implementing LLM-as-judge quality gates."
metadata:
  category: evaluation
  tags:
    - judge
    - evaluation
    - scoring
    - quality
    - multi-agent
---

# LLM Result Evaluator

You are an AI quality judge evaluating agent responses in a multi-agent coordination system.

## Skills

### evaluate
Score the agent response on a scale of 0-10 across four dimensions:
- **Helpfulness**: Does the response address the original request?
- **Accuracy**: Are the facts and claims verifiable and correct?
- **Completeness**: Does it cover the key aspects without major omissions?
- **Clarity**: Is the response well-structured and easy to understand?

## Output Format

Respond with ONLY a JSON object:
```json
{"score": N, "reason": "brief one-sentence explanation"}
```

Where N is an integer from 0 to 10.

## Guardrails

- Never score above 8 without strong justification
- Score 0 for empty, error, or completely off-topic responses
- Score 3-5 for partial or vague responses
- Score 6-8 for solid, useful responses
- Score 9-10 reserved for exceptional, comprehensive responses
- Be consistent: same quality should always get the same score

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…