Skip to content
Back to skills

Aeon Skill Evals

ASecurity

Validate the output of any installed skill against an assertion manifest — word counts, required patterns, forbidden phrases, required sections, source citation. Detects regressions by diffing vs prior runs (NEW_FAIL / NEW_PASS / CHRONIC / STABLE_FAIL). Bootstrap mode generates a starter manifest from a skill's recent successful runs so manifests aren't written speculatively. Triggers: "evaluate this skill's output", "check skill X for regressions", "bootstrap evals for Y", "did this skill ou...

  • 1,203 stars
  • 0 votes
  • 0 copies
  • 3 views
  • Added May 28, 2026
code-quality

Security analysis

A100/100

Scanned May 28, 2026

npx -y skills add BankrBot/skills --skill aeon-skill-evals --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Aeon Skill Evals?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Aeon Skill Evals
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/bankrbot-aeon-skill-evals/badge)](https://www.skillsdirectory.com/skills/bankrbot-aeon-skill-evals)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: aeon-skill-evals
description: |
  Validate the output of any installed skill against an assertion manifest — word counts, required
  patterns, forbidden phrases, required sections, source citation. Detects regressions by diffing
  vs prior runs (NEW_FAIL / NEW_PASS / CHRONIC / STABLE_FAIL). Bootstrap mode generates a starter
  manifest from a skill's recent successful runs so manifests aren't written speculatively.
  Triggers: "evaluate this skill's output", "check skill X for regressions", "bootstrap evals
  for Y", "did this skill output pass quality gates".
---

# aeon-skill-evals

Quality net for installed skills. Each skill can declare an assertion manifest; outputs are checked against it; failing assertions surface regressions and route concrete fixes.

## Manifest format

```yaml
token-movers:
  min_words: 200
  required_patterns: ["Top movers", "24h"]
  forbidden_patterns: ["I cannot", "as an AI"]
  must_cite_source: true
  min_distinct_items: 5

narrative-tracker:
  min_words: 400
  required_sections: ["TRANSITIONS", "POSITIONS", "MAP"]
  forbidden_patterns: ["exciting", "consider"]
  must_have_position_call: true
```

Supported assertions: `min_words` / `max_words`, `required_patterns` / `forbidden_patterns`, `required_sections`, `must_cite_source`, `min_distinct_items`, `output_pattern` (regex), and per-skill-family custom binary checks.

## Operations

- `eval` — run every manifest-defined skill against its latest output.
- `eval --skill=NAME` — one skill.
- `bootstrap --skill=NAME` — generate a starter manifest from recent successful runs.

## Regression states

| State | Action |
|---|---|
| `NEW_FAIL` | Passing last run, failing now. Severity scales with pass streak. |
| `NEW_PASS` | Failing last run, passing now. Log the win. |
| `CHRONIC` | Failing > 3 consecutive runs. Recommend operator review. |
| `STABLE_FAIL` | Always failing. Manifest assertion mismatch — flag for review. |

State in local `evals-state.json`.

## Bootstrap mode

Samples last 5 successful runs of a skill. Computes:

- `min_words` at p25 of historical runs.
- Required patterns from common section headers.
- Forbidden patterns from default list (refusals, hedging filler).

Emits the proposed manifest for review. Never auto-commits — assertions need a human signoff.

## Rules

- Assertions are observations, not specifications. Bootstrap before writing speculatively.
- Forbidden patterns catch hallucination markers and refusals. Keep the list tight; don't lint stylistic choices.
- Chronic failures get a recommendation, not a re-file.
- Manifest changes are reviewed; never auto-edited by this skill.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…