Skip to content
Back to skills

Os Skill Improvement

ASecurity

Continuously improves an existing agent skill based on eval results using the RED-GREEN-REFACTOR cycle. Apply when a skill's routing accuracy is low, trigger descriptions need sharpening, or os-eval-runner scores are below target. (1) run a RED baseline to observe the failure mode, (2) apply a focused patch and verify with os-eval-runner (GREEN), (3) refactor to close loopholes until score meets threshold. Integrates with os-eval-runner as the objective eval gate. NOT for scaffolding new skil...

  • 7 stars
  • 0 votes
  • 0 copies
  • 3 views
  • Added September 2, 2026
ai-agentspythonbashtestingrefactoringdocumentation

Security analysis

A100/100

Pro scans all 8 files and shows the line behind each finding

Scanned October 3, 2026

npx -y skills add richfrem/agent-plugins-skills --skill os-skill-improvement --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Os Skill Improvement?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Os Skill Improvement
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/richfrem-os-skill-improvement/badge)](https://www.skillsdirectory.com/skills/richfrem-os-skill-improvement)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: os-skill-improvement
version: 1.0.0
description: >
  Continuously improves an existing agent skill based on eval results using the
  RED-GREEN-REFACTOR cycle. Apply when a skill's routing accuracy is low, trigger
  descriptions need sharpening, or os-eval-runner scores are below target.
  (1) run a RED baseline to observe the failure mode,
  (2) apply a focused patch and verify with os-eval-runner (GREEN),
  (3) refactor to close loopholes until score meets threshold.
  Integrates with os-eval-runner as the objective eval gate.
  NOT for scaffolding new skills — use create-skill (agent-scaffolders) for that.
trigger: improve a skill, improve skill routing, fix routing accuracy, skill is not triggering,
  skill triggers too often, improve trigger description, update a skill trigger, skill patch,
  improve triggers, route a skill, routing precision, fix skill description, skill scoring low,
  eval score low, skill improvement, continuous skill improvement, refactor skill triggers,
  tdd for documentation, skill not routing correctly
allowed-tools: Read, Write, Edit, Bash
---

# Skill Improvement: RED-GREEN-REFACTOR (`os-skill-improvement`)

Adapts the RED-GREEN-REFACTOR cycle from software testing to skill authoring. A skill is a testable contract: always observe the failure BEFORE writing the fix.

## Contents

- [Critical Constraints](#critical-constraints)
- [Quick start](#quick-start)
- [Workflow](#workflow)
- [Verification](#verification)
- [References](#references)

## Critical Constraints

1. **Observe Failure First**: Never patch a skill without first running a RED scenario to observe empirical failure.
2. **Objective Eval Gate**: Every patch must be scored by `eval_runner.py` and achieve a KEEP verdict.
3. **No Scaffolding**: Use `create-skill` for scaffolding new skills; this skill only optimizes existing skills.

## Quick start

Snapshot the current skill evaluation baseline:

```bash
python3 scripts/eval_runner.py --skill <path/to/skill> --snapshot
```

## Workflow

1. **Observe RED Baseline**: Run pressure scenario without patch; document baseline failure mode.
2. **Declare Hypothesis**: State expected routing gain or precision fix before modifying files.
3. **GREEN Patch**: Edit `SKILL.md` description, trigger keywords, and representative examples.
4. **Evaluate**: Run `python3 scripts/eval_runner.py --skill <dir>`; require KEEP verdict.
5. **REFACTOR**: Close remaining edge cases and loopholes revealed by eval failures.

## Verification

Confirm skill passes evaluation with improved or equal score:

```bash
python3 scripts/eval_runner.py --skill <path/to/skill> --decision-only
```

## References

- [detailed-reference.md](references/detailed-reference.md) — TDD mapping, pressure scenarios, and baseline protocols.
- [acceptance-criteria.md](references/acceptance-criteria.md) — Acceptance criteria for continuous skill improvement.
- [fallback-tree.md](references/fallback-tree.md) — Remediation pathways when refactoring fails to converge.
- [skill_optimization_guide.md](references/operations/skill_optimization_guide.md) — Routing accuracy patterns and heuristics.
- [test-registry-protocol.md](references/testing/test-registry-protocol.md) — Test scenario registration standards.

Files in this skill

  • SKILL.md12.6 KB
  • evals/evals.json662 B
  • evals/results.tsv155 B
  • references/acceptance-criteria.md42 B
  • references/memory/improvement-ledger-spec.md56 B
  • references/operations/skill_optimization_guide.md61 B
  • references/testing/test-registry-protocol.md56 B
  • scripts/eval_runner.py31 B

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…