Skip to content
Back to skills

Refactor

ASecurity

Restructure or clean up code with no behavior change, proved by before-and-after checks. Use when: asked to clean up, extract, dedupe or simplify, even one function.

  • 446 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 2, 2026
ai-agentsgorefactoringgitapi

Works with

  • cli
  • api

Security analysis

A100/100

Pro scans all 3 files and shows the line behind each finding

Scanned October 5, 2026

npx -y skills add boshu2/agentops --skill refactor --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Refactor?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Refactor
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/boshu2-refactor-df2051cc/badge)](https://www.skillsdirectory.com/skills/boshu2-refactor-df2051cc)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: refactor
description: 'Restructure or clean up code with no behavior change, proved by before-and-after checks. Use when: asked to clean up, extract, dedupe or simplify, even one function.'
practices:
- refactoring
- legacy-code-seams
- design-patterns
hexagonal_role: supporting
consumes:
- repo-context
produces:
- code-changes
context_rel: []
skill_api_version: 1
user-invocable: true
context:
  window: fork
  intent:
    mode: task
  sections:
    exclude:
    - HISTORY
metadata:
  capabilities: [refactor]
  effects: [modify_source_files]
  canonical_status: canonical
  disposition: keep_specialist
  tier: execution
  dependencies: []
output_contract: code changes with regression evidence
---
# Refactor — one structural experiment

Refactor changes structure while preserving observable behavior. It performs one
caller-selected transformation and reports the result. "Behavior-preserving" is
a claim to prove with before/after checks, never to assert.

## What counts as behavior

Unless the caller explicitly excluded a surface, all of these must survive:

- **Messages and exit codes.** Error and output text compares byte-for-byte;
  exit codes, error types and which input raises which error stay the same.
  Scripts and callers parse them. Preserve an inconsistent message and report
  it; normalizing it is a behavior change.
- **Differences between near-duplicates.** When merging duplicated branches,
  carry every difference (constants, comparisons, messages, extra steps) as a
  parameter or a branch. Do not unify a difference the caller has not declared
  accidental.
- **Interfaces.** Public signatures, defaults, return types, persisted field
  names, protocol values and CLI flags. Renaming one is a compatibility change
  unless the accepted scope provides for it.
- **Order and coverage.** Branch priority, default handling, evaluation count,
  side-effect order, and the set of tests that run. A pre-existing red that
  vanishes, or a test that stops running, is a behavior change.

## Procedure

1. Name the preserved behavior, the focused acceptance surface and the concrete
   structural problem for its callers, in the caller's domain terms.
2. Run the focused check and the smallest regression check the changed surface
   justifies, and record that honest baseline, including reproducible ambient
   failures. For an evaluation comparing executable behavior, pin the starting
   source, build its baseline before edits and keep that binary and the
   comparison inputs.
3. Apply one bounded transformation: extract, rename, inline, simplify,
   encapsulate, move, or delete dead code. Judge it by what callers must
   understand and where a domain rule must change, not by file size.
4. A bug or suspicious inconsistency found on the way is reported separately
   (location, why it looks wrong) and left unfixed. Fixing it inside the
   refactor hides a behavior change the caller did not authorize.
5. Rerun the same focused check and the smallest justified regression check
   over the same inputs, including error paths.
6. Report, then stop. A red result is evidence for the caller; this skill does
   not revert, narrow, retry, commit, validate, or route subsequent work.

When nothing can be executed (no runtime, no tests, code supplied in a
message), neutrality is unproven: give the exact before/after commands and
inputs the caller must run, error paths included, and list every surface under
behavior not checked.

```text
transformation: <the one change>
preserved:      <behavior and surfaces from step 1>
checks:         <command>: before -> <result>; after -> <result>   (or "not run")
outputs:        <before/after hashes when the surface produces output>
diff:           <files touched>; only those the transformation names
suspected bugs: <file:line, why>; reported, not fixed
not checked:    <surfaces no check covered>; present even when empty
```

## Responsibility and interface cost

Before adding an interface or splitting a module, inspect representative callers.
Count the concepts they must coordinate: required setup, ordering, states, error
handling and repeated domain rules. A useful boundary puts a cohesive rule under
one owner and lets callers request an outcome without reproducing that rule.
Reject a wrapper that only adds another name or pushes the same coordination
into its callers. Existing boundaries are sufficient when no concrete caller
problem warrants changing them.

Use the caller's vocabulary for extracted operations and types. A naming
ambiguity that changes behavior belongs with the existing domain definition;
consult [Domain](../domain/SKILL.md) only when that distinction needs work.

When the transformation needs a seam — an extraction boundary, interface, or
module split — and more than one candidate seam exists, probe before you cut.
Run the probe in disposable isolation (a scratch branch, worktree, or copied
tree the caller's policy allows): rough in the seam, see what it forces —
signature churn, import cycles, test rewrites — then discard the probe and
keep only the knowledge. Stop condition: at most two probes; if the second
candidate seam also fights back, report both findings to the caller instead of
trying a third. Cutting the first imaginable seam directly into the working
tree is the **premature seam** failure mode: the wrong boundary calcifies
because reverting it now costs more than living with it.

## Neutrality gates

Gate the transformation on behavior-identical proof:

- The focused check and the package-level regression check pass both before
  and after, with the same set of pre-existing failures: no new red and no
  vanished red.
- For output-producing surfaces (generators, serializers, formatters, reports),
  capture output hashes over identical inputs before the change and compare
  byte-for-byte after. A mismatch is a behavior diff to surface and explain,
  never to shrug at; the caller decides whether to keep, narrow, or reverse it.

A neutrality gate that was skipped or narrowed after the fact is the
**post-hoc neutrality** failure mode — the diff decides what got tested. Name
any surface the gates did not cover under behavior not checked.

## References

- [Behavior-preserving simplification](references/behavior-preserving-simplification.md) — refactoring catalog and per-pattern safety checks
- [Behavior scenarios](references/refactor.feature)
- [Upstream capability reference](https://github.com/mattpocock/skills/blob/main/skills/engineering/codebase-design/SKILL.md) — Matt Pocock; original AgentOps adaptation.

Files in this skill

  • SKILL.md3.5 KB
  • references/behavior-preserving-simplification.md5.2 KB
  • references/refactor.feature2 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…