Skip to content
Back to skills

Evaluate Plugin

ASecurity

Evaluate a local Codex plugin in engineer-friendly language. Use when the user says "evaluate this plugin", "audit this plugin", "why did this score that way", "what should I fix first", "help me benchmark this plugin", or asks for a plugin-wide report before comparing versions.

  • 8 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 12, 2026
toolsbash

Security analysis

A100/100

Pro scans all 8 files and shows the line behind each finding

Scanned September 12, 2026

npx -y skills add stanfish06/skillquarium --skill evaluate-plugin --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Evaluate Plugin?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Evaluate Plugin
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/stanfish06-evaluate-plugin/badge)](https://www.skillsdirectory.com/skills/stanfish06-evaluate-plugin)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: evaluate-plugin
description: Evaluate a local Codex plugin in engineer-friendly language. Use when the user says "evaluate this plugin", "audit this plugin", "why did this score that way", "what should I fix first", "help me benchmark this plugin", or asks for a plugin-wide report before comparing versions.
---

# Evaluate Plugin

Use this skill when the target is a plugin root with `.codex-plugin/plugin.json`.

## Workflow

1. Treat "Evaluate this plugin." as the default entrypoint.
2. If the request comes in as natural chat language, use `plugin-eval start <plugin-root> --request "<user request>" --format markdown` first so the user sees the routed local path.
3. Run `plugin-eval analyze <plugin-root> --format markdown`.
4. Read `Fix First` before drilling into manifest findings, nested skill findings, and code or coverage details.
5. If the plugin contains multiple skills, summarize the strongest and weakest ones explicitly.
6. If the user wants measured usage, switch to "Help me benchmark this plugin." and use the starter benchmark flow.
7. If the user wants trend data, compare two JSON outputs with `plugin-eval compare`.

## Chat Requests To Recognize

- `Evaluate this plugin.`
- `Audit this plugin.`
- `Why did this score that way?`
- `What should I fix first?`
- `Help me benchmark this plugin.`
- `What should I run next?`

## Commands

```bash
plugin-eval start <plugin-root> --request "Evaluate this plugin." --format markdown
plugin-eval analyze <plugin-root> --format markdown
plugin-eval start <plugin-root> --request "What should I run next?" --format markdown
plugin-eval compare before.json after.json
plugin-eval report result.json --format html --output ./plugin-eval-report.html
plugin-eval init-benchmark <plugin-root>
plugin-eval benchmark <plugin-root> --dry-run
```

## Reference

- `references/chat-first-workflows.md`

Files in this skill

  • SKILL.md1.8 KB
  • agents/openai.yaml98 B
  • references/benchmark-harness.md1.4 KB
  • references/chat-first-workflows.md3 KB
  • references/evaluation-result-schema.md2.8 KB
  • references/metric-pack-manifest.md1.3 KB
  • references/observed-usage.md1.3 KB
  • references/technical-design.md10.6 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…