Skip to content
Back to skills

Zero Shot Prompt Robust Evaluation

ASecurity

Compute held-out prompted-task accuracy and prompt wording consistency for T0/P3-style recovery experiments.

  • 247 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 9, 2026
testing

Security analysis

A100/100

Scanned September 9, 2026

npx -y skills add VectorSpaceLab/AREX-Skill --skill zero_shot_prompt_robust_evaluation --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Zero Shot Prompt Robust Evaluation?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Zero Shot Prompt Robust Evaluation
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vectorspacelab-zero-shot-prompt-robust-evaluation/badge)](https://www.skillsdirectory.com/skills/vectorspacelab-zero-shot-prompt-robust-evaluation)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

SKILL.md
---
name: zero_shot_prompt_robust_evaluation
description: Compute held-out prompted-task accuracy and prompt wording consistency for T0/P3-style recovery experiments.
---

# Zero-Shot and Prompt-Robust Evaluation

Use this skill after a model or proxy predictor has produced predictions for held-out prompted examples. It measures correctness and robustness across alternative templates for the same raw item.

## Inputs
- Prediction records with `example_id`, `template_id`, `prediction`, and `target`.

## Outputs
- Overall accuracy, per-template accuracy, and same-example prompt consistency.

## Workflow
1. Compare canonical prediction and target strings for exact-match accuracy.
2. Group records by template id for per-prompt diagnostics.
3. Group records by example id and measure whether all prompt variants produce the same prediction.
4. Return JSON-serializable metrics and counts.

## Validation
Run included deterministic metric tests.

## Limitations
This skill does not generate predictions; it only evaluates records already produced by an experiment.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…