Skip to content
Back to skills

Os Eval Runner

ASecurity

Stateless evaluation engine that scores and gates skill improvement iterations using headless Python evaluation scripts. Use when the user says "evaluate this skill", "run autoresearch loop on", "optimize this skill", "run the eval loop", or when another agent proposes a change and needs validation.

  • 7 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 2, 2026
ai-agentspythonbash

Security analysis

A100/100

Pro scans all 19 files and shows the line behind each finding

Scanned October 3, 2026

npx -y skills add richfrem/agent-plugins-skills --skill os-eval-runner --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Os Eval Runner?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Os Eval Runner
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/richfrem-os-eval-runner/badge)](https://www.skillsdirectory.com/skills/richfrem-os-eval-runner)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: os-eval-runner
plugin: agent-agentic-os
description: >
  Stateless evaluation engine that scores and gates skill improvement iterations using
  headless Python evaluation scripts. Use when the user says "evaluate this skill",
  "run autoresearch loop on", "optimize this skill", "run the eval loop", or when
  another agent proposes a change and needs validation.
allowed-tools: Read, Write, Edit, Bash, Glob, Grep
---

# Skill Improvement Evaluator (`os-eval-runner`)

Stateless evaluation engine that scores and gates skill improvement iterations using headless Python evaluation scripts.

## Contents

- [Critical Constraints](#critical-constraints)
- [Quick start](#quick-start)
- [Workflow](#workflow)
- [Verification](#verification)
- [References](#references)

## Critical Constraints

1. **Ownership Boundary**: Engine owns scripts (`evaluate.py`, `eval_runner.py`) and templates. Experiment state (`references/program.md`, `evals/evals.json`, `evals/results.tsv`) deploys strictly alongside target.
2. **Objective Verification Invariant**: Never mentally simulate routing or score subjective impressions. All evaluations must be executed programmatically via headless python scripts.
3. **Holdout & Overfitting Gate**: Loops require a locked holdout prompt set. If holdout score drops, the candidate mutation must be forcefully discarded.

## Quick start

Establish an initial baseline evaluation for a target skill:

```bash
python3 scripts/evaluate.py --skill <path/to/skill> --baseline --desc "initial baseline"
```

## Workflow

1. **Intake & Discovery**: Confirm target skill path, metric to optimize (`quality_score`, `f1`, etc.), and operating mode (Mode 1: Autoresearch Loop vs Mode 2: Single-shot QA).
2. **Scaffold Experiment State**: Ensure `evals.json` and `program.md` exist alongside target; scaffold via `scripts/init_autoresearch.py` if missing.
3. **Execution Mode**:
   - **Mode 1 (Autoresearch Loop)**: Iteratively request mutations via proposer, run `evaluate.py`, and keep or discard based on score deltas.
   - **Mode 2 (Single-shot QA)**: Evaluate proposed candidate against baseline to decide KEEP (`exit 0`) or DISCARD (`exit 1`, revert).
4. **Overfitting Check**: Run holdout set evaluation before accepting any candidate improvement.

## Verification

Execute an end-to-end smoke test verifying script execution and exit codes:

```bash
python3 scripts/init_autoresearch.py --experiment-dir temp/test-exp --mutation-target SKILL.md
python3 scripts/evaluate.py --skill temp/test-exp --baseline --desc "smoke test"
```

## References

- [quickstart-setup.md](references/quickstart-setup.md) — 4-step setup and re-baselining procedure.
- [mode-1-loop-protocol.md](references/mode-1-loop-protocol.md) — Autoresearch loop protocol and mutation cycles.
- [mode-2-qa-protocol.md](references/mode-2-qa-protocol.md) — Single-shot QA diff validation.
- [overfitting-gate.md](references/overfitting-gate.md) — Holdout set checks and forced discard logic.
- [cheapest_models.md](references/cheapest_models.md) — Recommended lightweight models for evaluation runs.
- [post_run_survey.md](references/memory/post_run_survey.md) — Mandatory post-run survey template.
- [survey-protocol.md](references/survey-protocol.md) — Evaluator survey and retrospective guidelines.

Files in this skill

  • QUICKSTART.md4.2 KB
  • README.md3 KB
  • SKILL.md3.9 KB
  • assets/templates/autoresearch/copilot_proposer_prompt.md.template80 B
  • assets/templates/autoresearch/evals.json.template64 B
  • assets/templates/autoresearch/program.md.template64 B
  • assets/templates/autoresearch/results.tsv.template65 B
  • assets/templates/eval-instructions.template.md58 B
  • evals.json759 B
  • evals/evals.json1.3 KB
  • evals/results.tsv1.7 KB
  • references/acceptance-criteria.md42 B
  • references/autoresearch-architecture.md48 B
  • references/autoresearch-overview.md44 B
  • references/autoresearch-program-md-overview.md55 B
  • references/cheapest_models.md38 B
  • references/diagrams/autoresearch-loop.mmd53 B
  • references/diagrams/mapping-karpathy-to-skill-improvement-eval.mmd78 B
  • references/lab-space-protocol.md41 B

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…