Skip to content
Back to skills

Mmlu Item Schema

ASecurity

Validate and canonicalize MMLU-style four-option subject examples for evaluation workflows.

  • 247 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 9, 2026
toolspython

Security analysis

A100/100

Scanned September 9, 2026

npx -y skills add VectorSpaceLab/AREX-Skill --skill mmlu_item_schema --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Mmlu Item Schema?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Mmlu Item Schema
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vectorspacelab-mmlu-item-schema/badge)](https://www.skillsdirectory.com/skills/vectorspacelab-mmlu-item-schema)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

SKILL.md
---
name: mmlu_item_schema
description: Validate and canonicalize MMLU-style four-option subject examples for evaluation workflows.
---

# mmlu_item_schema

Use this skill when a recovery or evaluation task needs the MMLU mechanism represented by this module. Do not use it as a substitute for a full model evaluation unless the run explicitly declares a reduced proxy target.

## Inputs
Inputs are JSON-compatible dictionaries following the module document contract. Required fields are validated by the script in `scripts/`.

## Outputs
Outputs are JSON-compatible dictionaries with normalized labels, prompts, scores, or recovery artifacts. Downstream modules should consume these outputs instead of parsing free-form text.

## Workflow
1. Read the paper profile and module plan to confirm the selected target.
2. Call the script or import the helper functions from `scripts/`.
3. Preserve the boundary between item construction, prompt construction, scoring, and recovery logging.
4. Run the validation command after changes: `python <distiller>/module-to-skill/scripts/validate_skill_tree.py <skill_dir> --run-tests`.

## Limitations
This skill captures the evaluation mechanism, not the full hosted GPT-3 runtime or the full private benchmark distribution. Soft-mode recovery must label reduced results as proxy evidence.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…