Skip to content
Back to skills

Advanced Evaluation

ASecurity

Production-grade techniques for evaluating LLM outputs using LLM-as-judge approaches with bias mitigation.

  • 7 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 3, 2026
ai-agentsapi

Works with

  • api

Security analysis

A100/100

Pro scans all 2 files and shows the line behind each finding

Scanned September 3, 2026

npx -y skills add mediar-ai/skillhubz --skill advanced-evaluation --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Advanced Evaluation?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Advanced Evaluation
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/mediar-ai-advanced-evaluation/badge)](https://www.skillsdirectory.com/skills/mediar-ai-advanced-evaluation)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
# Advanced Evaluation

Production-grade techniques for evaluating LLM outputs using LLM-as-judge approaches with bias mitigation.

## Prerequisites

- Understanding of evaluation metrics
- Access to LLM APIs for judge models

## Instructions

### Core Approaches

**Direct Scoring**: Single LLM rates one response on a defined scale.
- Best for: Objective criteria (factual accuracy, instruction following)
- Requires: Clear criteria, calibrated scale, chain-of-thought justification

**Pairwise Comparison**: LLM compares two responses and selects the better one.
- Best for: Subjective preferences (tone, style, persuasiveness)
- Requires: Position bias mitigation (swap positions and check consistency)

### Bias Mitigation

| Bias | Mitigation |
|------|------------|
| Position Bias | Evaluate twice with swapped positions |
| Length Bias | Explicit prompting to ignore length |
| Self-Enhancement | Use different models for generation and evaluation |
| Verbosity Bias | Criteria-specific rubrics |

### Pairwise Comparison Protocol

1. First pass: Response A first, Response B second
2. Second pass: Response B first, Response A second
3. Consistency check: If passes disagree, return TIE
4. Final verdict: Consistent winner with averaged confidence

### Rubric Components

1. **Level descriptions**: Clear boundaries for each score
2. **Characteristics**: Observable features per level
3. **Examples**: Representative text (optional but valuable)
4. **Edge cases**: Guidance for ambiguous situations

### Decision Framework

```
Is there objective ground truth?
├── Yes → Direct Scoring (factual accuracy, format compliance)
└── No → Is it preference/quality judgment?
    ├── Yes → Pairwise Comparison (tone, creativity)
    └── No → Reference-based evaluation
```

## Guidelines

1. Always require justification before scores (15-25% reliability improvement)
2. Always swap positions in pairwise comparison
3. Match scale granularity to rubric specificity
4. Separate objective and subjective criteria
5. Include confidence scores calibrated to evidence strength

## Notes

- Chain-of-thought prompting improves evaluation reliability
- Single-pass pairwise comparison is corrupted by position bias
- Validate automated evaluation against human judgments

Source: muratcankoylan/Agent-Skills-for-Context-Engineering

Files in this skill

  • manifest.json354 B
  • skill.md2.3 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…