Skip to content
Back to skills

Test Reviewer

ASecurity

Review existing tests against the software and contracts they actually exercise. Use to assess whether a test has justified unique protection, detects consequential faults, tolerates valid refactors, earns its lifecycle cost, or deserves strengthening, replacement, consolidation, or removal consideration. Reviews are read-only and evidence-backed, including TDD and regression tests.

  • 67 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 29, 2026
testinggotestingsecuritydocumentation

Security analysis

A100/100

Pro scans all 9 files and shows the line behind each finding

Scanned October 6, 2026

npx -y skills add Jamie-BitFlight/claude_skills --skill test-reviewer --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Test Reviewer?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Test Reviewer
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/jamie-bitflight-test-reviewer/badge)](https://www.skillsdirectory.com/skills/jamie-bitflight-test-reviewer)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: test-reviewer
description: Review existing tests against the software and contracts they actually exercise. Use to assess whether a test has justified unique protection, detects consequential faults, tolerates valid refactors, earns its lifecycle cost, or deserves strengthening, replacement, consolidation, or removal consideration. Reviews are read-only and evidence-backed, including TDD and regression tests.
---

# Test Reviewer

Read [Testing principles](../../docs/testing-principles.md) before review; it defines the shared
criteria, compact test card, safe execution boundary, and evidence required for each disposition.

## Scope

Review without editing or deleting the target's tests, product, requirements, or generated state.
A request to review is not permission to execute destructive probes or update snapshots. Use safe
isolated execution when authorized and available; otherwise report the proposed check as unvalidated.
Prefer a reviewer independent of the test author where available; disclose when that is not possible.

## Procedure

1. **Resolve architecture, contract altitude, and the review boundary.** Read the governing
   repository instructions and architecture needed for the scope. Trace important tests upward from
   their observation boundary to the stable component/public/system contract or holistic goal they
   protect. Identify executable components/responsibilities, stable
   user/public/machine/cross-component contracts, destructive or durable-state boundaries,
   authorization/security/data-integrity risks, and cheaper deterministic enforcement from typing,
   schemas, compilation, linting, static analysis, or contract validators. Inventory the scoped
   tests, fixtures, production entry points/consumers, and CI selection, then map tests into
   provisional protection families by claim, boundary, and relevant failure mechanism. Read only
   the dependencies needed to decide a disposition; carry remaining scope as an explicit gap.
2. **Trace at the resolution that can change the decision.** Deep-trace high-consequence,
   suspicious, high-cost, or uncertain families first. For each disposition-relevant test or
   representative family member, map setup -> production path -> observation -> assertion.
   Identify its intended claim and authority, the failure consequence, what it actually observes,
   and what can remain broken while it passes. Split a family whenever variants, boundaries,
   forbidden effects, diagnostics, or fault mechanisms differ materially. Separate characterization
   from approved intent. Do not infer importance from names, test counts, or covered lines.
3. **Challenge effectiveness.** Apply the shared oracle, fault-sensitivity, refactor-tolerance,
   fidelity, isolation, and diagnostic criteria. Prefer evidence from stable
   system/integration/contract/lifecycle behavior over implementation details when both address the
   same risk. Select relevant controls only when they materially improve confidence. Confirm any
   original defect or seeded fault reaches the intended observation, not collection/setup failure.
   Do not require mutation execution for every test or call static inspection observed fault
   detection.
4. **Assess maintenance-adjusted value.** Apply the shared economic model to each
   disposition-relevant test/family. Record expected protection benefit separately from expected
   ownership cost. Benefit includes consequence avoided, regression exposure, detection
   effectiveness, unique protection, contract durability, and feedback value. Ownership cost
   includes human/agent context, implementation/prose coupling and churn, fixtures/data, execution,
   diagnosis, flakiness, environment/dependencies, refactor drag, duplication, and
   review/coordination. Compare tests by the consequential failures they uniquely exclude and the contract altitude
   they protect, not by pyramid quotas. Treat unit tests that fail under behavior-preserving refactors
   because private helpers, constants, branches, or decomposition changed as implementation-coupled
   unless those details are authoritative contracts. Prefer a broader lifecycle/contract carrier
   when it provides equivalent or stronger useful protection with acceptable diagnostics.
5. **Recommend a disposition.** Use KEEP, STRENGTHEN, REPLACE, CONSOLIDATE, REMOVAL-CANDIDATE,
   or UNRESOLVED with the evidence defined in the shared guide. Do not use missing documentation,
   high cost, or unmeasured sensitivity alone as proof of uselessness. For consolidation/removal,
   answer: if this test disappeared, which plausible important regression could now pass undetected?
   Identify the surviving carrier for every required guarantee and any residual risk.
6. **Specify the next validation.** Distinguish product, requirement, oracle, boundary/interface,
   and test/CI-harness defects. For a proposed change, define success and protected behavior before
   its implementation; compare baseline and candidate under equivalent relevant conditions, adding
   independent/unseen scenarios for substantial redesign where practical. No corrections are
   applied by this review. Carry unresolved authority or evidence into the handoff.

## Report

Return scope/revision, sources read, and uncovered scope before the findings. For each
disposition-relevant test or explicitly bounded protection family, use a row containing:

```text
Test/location | intended claim + authority | actual observation/boundary
Protection benefit | ownership cost | effectiveness + evidence | disposition
Correction category | surviving protection / gap | next discriminating validation
```

Distinguish observed defects, inferred risks, and proposed probes. Report claim -> validator ->
observed result -> evidence boundary. Include the exact command/selection when run, outcomes,
skips/xfails, omitted suites, and unavailable environments. Authored evals and source inspection
are not passing runtime evidence.

Conclude with `REVIEW COMPLETE` when the declared scope was assessed, or `REVIEW PARTIAL` with
unread dependencies or missing evidence; separately state whether required behavior is supported,
failed, or unvalidated. Report `No material test findings in the assessed scope` when applicable.
Completion of a review never means every test passed, and a removal candidate never authorizes deletion.

For findings needing causal investigation, read
[Test Failure Mindset](../test-failure-mindset/SKILL.md) or
[Root-Cause Tracing Process](../root-cause-tracing-process/SKILL.md); carry the contract and evidence
forward. For new/replacement test planning use [Test Designer](../test-designer/SKILL.md).
When project commands are unknown, [DH Meta Docs](../dh-meta-docs/SKILL.md) routes quality-gate
discovery. Preserve the caller's existing SAM verdict, artifact registration, and status contract.

Files in this skill

  • SKILL-GOALS.md278 B
  • SKILL.md5 KB
  • evals/README.md309 B
  • evals/evals.json6.1 KB
  • evals/fixtures/review_project/CONTRACT.md975 B
  • evals/fixtures/review_project/__init__.py87 B
  • evals/fixtures/review_project/app.py1 KB
  • evals/fixtures/review_project/pytest.ini64 B
  • evals/fixtures/review_project/test_app.py872 B

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…