Skip to content
Back to skills

Testing And Qa

ASecurity

Test and evaluate Agent Skill performance with benchmarks

  • 3 stars
  • 0 votes
  • 0 copies
  • 3 views
  • Added September 12, 2026
ai-agentsgotestinggitperformance

Security analysis

A100/100

Pro scans all 21 files and shows the line behind each finding

Scanned September 12, 2026

npx -y skills add thedixitjain/the-mega-skill-library --skill testing-and-qa --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Testing And Qa?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Testing And Qa
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/thedixitjain-testing-and-qa/badge)](https://www.skillsdirectory.com/skills/thedixitjain-testing-and-qa)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: test-skill
description: "Test and evaluate Agent Skill performance with benchmarks"
category: testing-and-qa
source_repo: libukai/awesome-agent-skills
source_path: "plugins/agent-skills-toolkit/1.0.0/commands/test-skill.md"
source_url: https://github.com/libukai/awesome-agent-skills/blob/HEAD/plugins/agent-skills-toolkit/1.0.0/commands/test-skill.md
---


# Test and Evaluate Skill

You are helping the user test and evaluate an Agent Skill's performance.

**IMPORTANT**: First invoke `/agent-skills-toolkit:skill-creator-pro` to load the complete testing and evaluation framework, including scripts and evaluation tools.

Once skill-creator-pro is loaded, use the evaluation workflow and tools:

## Quick Testing Process

1. **Prepare Test Cases**
   - Review existing test prompts
   - Add new test cases if needed
   - Cover various scenarios

2. **Run Tests** (use skill-creator-pro's scripts)
   - Execute test prompts with the skill
   - Use `scripts/run_eval.py` for automated testing
   - Use `scripts/run_loop.py` for batch testing
   - Collect results and outputs

3. **Qualitative Evaluation**
   - Review outputs with the user
   - Use `eval-viewer/generate_review.py` to visualize results
   - Assess quality and accuracy
   - Identify improvement areas

4. **Quantitative Metrics** (use skill-creator-pro's tools)
   - Run `scripts/aggregate_benchmark.py` for metrics
   - Measure success rates
   - Calculate variance analysis
   - Compare with baseline

5. **Generate Report**
   - Use `scripts/generate_report.py` for comprehensive reports
   - Summarize test results
   - Highlight strengths and weaknesses
   - Provide actionable recommendations

## Available Tools from skill-creator-pro

- `scripts/run_eval.py` - Run evaluations
- `scripts/run_loop.py` - Batch testing
- `scripts/aggregate_benchmark.py` - Aggregate metrics
- `scripts/generate_report.py` - Generate reports
- `eval-viewer/generate_review.py` - Visualize results
- `agents/grader.md` - Grading subagent
- `agents/analyzer.md` - Analysis subagent
- `agents/comparator.md` - Comparison subagent

## Evaluation Criteria

- **Accuracy**: Does it produce correct results?
- **Consistency**: Are results reliable across runs?
- **Completeness**: Does it handle all use cases?
- **Efficiency**: Is the workflow optimal?
- **Usability**: Is it easy to trigger and use?

## Next Steps

Based on test results:
- Run `/agent-skills-toolkit:improve-skill` to address issues
- Expand test coverage for edge cases
- Document findings for future reference

---

**Source:** [`libukai/awesome-agent-skills`](https://github.com/libukai/awesome-agent-skills) → `plugins/agent-skills-toolkit/1.0.0/commands/test-skill.md`

**Also appears in:** `libukai/awesome-agent-skills/plugins/agent-skills-toolkit/1.1.0/commands/test-skill.md`, `libukai/awesome-agent-skills/plugins/agent-skills-toolkit/1.2.0/commands/test-skill.md`

Files in this skill

  • ab-testing-anti-patterns.md9.5 KB
  • ab-testing-decisions.md9.9 KB
  • ab-testing-patterns.md8.4 KB
  • ab-testing-sharp-edges.md9.1 KB
  • add-mutation-testing.md4.1 KB
  • add-property-based-testing.md4.3 KB
  • analyze-coverage.md2.2 KB
  • analyze-failures.md1.2 KB
  • analyze-test-failures.md4.7 KB
  • analyze-tests-command.md2.1 KB
  • anti-patterns-qa-engineering.md16.9 KB
  • assess-bug.md9.8 KB
  • automated-unit-test-generation.md10.7 KB
  • bossqa.md2.6 KB
  • browser-test.md5.2 KB
  • bug-fix.md898 B
  • bug-report-assistant.md5.7 KB
  • capture-baseline.md1.9 KB
  • comando-e2e--affaan-m.md11.2 KB
  • comando-e2e.md11.3 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…