Skip to content
Back to skills

Skill Compressor

ASecurity

Make an agent skill cheaper to load without breaking when it activates or how it behaves. Use when a skill has accumulated too much text and the smaller version must earn promotion through routing and runtime tests.

  • 5 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 19, 2026
ai-agents

Security analysis

A100/100

Pro scans all 6 files and shows the line behind each finding

Scanned September 19, 2026

npx -y skills add simonasrazm/skills --skill skill-compressor --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Skill Compressor?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Skill Compressor
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/simonasrazm-skill-compressor/badge)](https://www.skillsdirectory.com/skills/simonasrazm-skill-compressor)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: skill-compressor
description: Make an agent skill cheaper to load without breaking when it activates or how it behaves. Use when a skill has accumulated too much text and the smaller version must earn promotion through routing and runtime tests.
---

# Skill Compressor

Optimize behavior per loaded token, not file length. A smaller candidate wins only
when it preserves or improves every required quality cell. Treat plausible wording
changes as hypotheses until execution evidence supports them.

## Compression loop

1. **Freeze:** preserve the exact baseline, promotion bank, graders, settings,
   thresholds, staged evidence budget, and stop rules. Include real failures.
2. **Measure:** inventory description, always-loaded body, each routed reference,
   full surface, and expected loaded tokens. Run `scripts/measure_skill.py` with
   `--require-tokenizer` or use an available tokenizer; record encoding and version.
   Label an unverified target-model mapping. Use provider receipts for execution
   input, cached input, output and reasoning; keep subsets separate. Character or
   word counts cannot qualify a token-saving candidate. Provide observed route frequencies.
3. **Map:** create a behavioral ledger: stable rule ID, decision changed, activation
   condition, owning file, dependent rules, positive case, and failure if lost.
   Separate catalog routing from post-load execution behavior.
4. **Hypothesize:** consider unchanged control, removal, and the shortest replacement
   per seam: concise semantic wording, established pattern names with necessary
   qualifiers, and telegraphic notation. Measure each against the unchanged text;
   symbols and abbreviations are not inherently fewer tokens. Prefer one independent
   variable. Reject noncompetitive variants by
   inspection; generating hypotheses does not require executing them.
5. **Screen:** run deterministic checks, then one observation per live variant on the
   most discriminating known failure. Reuse a condition-identical control observation.
   One hard regression may reject; one clean screen cannot promote.
6. **Falsify:** before replication, challenge survivors on a different failure
   mechanism, archetype, or grader. Run another trial only when its result can change
   the decision. Localize failures by restoring or splitting the changed seam.
7. **Transfer:** replicate survivors on the frozen required cells, including fresh
   held-out and compounded cases. Promote only when every cell passes and no material
   NFR regresses. Otherwise keep the baseline and report the smallest failing seam.
8. **Apply:** update the source, validate it, and prove it is byte-identical to the
   accepted candidate. Rerun the bank only if application transformed the bytes;
   report baseline, candidate, and installed hashes plus measured deltas.

Read [evaluation protocol](references/evaluation-protocol.md) before designing the
test bank. Read [transformation mechanics](references/transformation-mechanics.md)
when classifying or rewriting seams.

## Optimization rules

- Optimize `catalog + always-loaded + routed-on-demand` expected cost. Moving text
  to a reference is not compression when that reference always loads.
- Preserve behavioral atoms, not necessarily their original sentences. A redundant,
  obsolete, default, contradictory, or harmful atom may be removed only by ablation.
- Test both activation and restraint: cases where guidance is needed, irrelevant,
  misleading, and accumulated across multiple turns or deliverables.
- Do not tune to evaluator nouns, fixtures, exact phrases, or one domain. Each retained
  mechanism must generalize to multiple archetypes or an explicit invariant.
- Do not accept aggregate wins that hide a failed task cell or catastrophe.
- Model-authored time estimates are not NFR evidence; use executor timestamps.
- Keep raw prompts, outputs, traces, grades, invalid runs, and protocol deviations.
- Stop when remaining candidates either fail quality gates or save less than the
  preregistered minimum practical token delta.

Files in this skill

  • SKILL.md4 KB
  • agents/openai.yaml116 B
  • references/evaluation-protocol.md4.6 KB
  • references/transformation-mechanics.md2.3 KB
  • scripts/measure_skill.py5.6 KB
  • scripts/test_measure_skill.py3.1 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…