Skip to content
Back to skills

Plan Exec

ASecurity

Use whenever the user asks to execute, implement, resume or continue a written plan — a `.context/plans/` document or any plan with checkboxes/phases — however small the phase looks, even a single-file edit: the plan''s checkboxes and Execution log are updated by this skill, so do not carry out the phase directly from the plan text. Fires on "implement the plan", "execute plan X", "let''s execute the plan", "continue with phase Y", "resume the plan", "run phase 1 of the plan", "run the plan p...

  • 2 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 19, 2026
ai-agentsgobashtestingcode-reviewgit

Security analysis

A100/100

Pro scans all 16 files and shows the line behind each finding

Scanned October 1, 2026

npx -y skills add yacb2/aidex --skill plan-exec --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Plan Exec?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Plan Exec
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/yacb2-plan-exec/badge)](https://www.skillsdirectory.com/skills/yacb2-plan-exec)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: plan-exec
description: 'Use whenever the user asks to execute, implement, resume or continue a written plan — a `.context/plans/` document or any plan with checkboxes/phases — however small the phase looks, even a single-file edit: the plan''s checkboxes and Execution log are updated by this skill, so do not carry out the phase directly from the plan text. Fires on "implement the plan", "execute plan X", "let''s execute the plan", "continue with phase Y", "resume the plan", "run phase 1 of the plan", "run the plan phase by phase". Enforces between-phase discipline: code-review, commit, handoff when context grows. Not for: creating the plan itself (/aidex:plan); one-shot tasks with no phases; bug fixes (/aidex:bugfix); pure refactors with no plan document.'
disable-model-invocation: false
allowed-tools: Bash Read Write Edit Agent
model-policy: per-stage
---

# Plan Execution

Drive the implementation of a written multi-phase plan with consistent
between-phase discipline: review the diff, commit, and hand off the session
when context grows. This skill centralizes
the workflow so the user does not have to repeat it in every prompt.

## Default autonomy

On run start, apply [Mode A autonomy](../conventions/references/autonomy-conventions.md)
automatically — do not wait for the user to grant it. Questions live in the
initial alignment moment only; after that the run proceeds start-to-finish
(deny/pre-authorized/mandated/autonomous — see "Operating mode" below).

## Operating mode

**Front-loaded, then autonomous start-to-finish.** Resolve every question at
**Orient** (phase 0); after that, run all phases without interrupting. Follow the
shared autonomy canon
([autonomy-conventions.md](../conventions/references/autonomy-conventions.md)).
The operative rule here:

- **Ask everything up front, at Orient.** Surface clarifications and confirm any
  publication the plan implies (deploy/publish/release) before phase 1. If the plan
  did not pre-authorize a publish step, surface it at the **end** — not mid-run.
- **Evaluate batch-promotion at Orient (mandatory, one line).** Before phase 1, decide
  whether the plan's `afk-impl` phases should run as a durable `Workflow`, and say so in
  one line. The full rule is in
  [`references/01-unattended-batch-execution.md`](references/01-unattended-batch-execution.md)
  § Promotion at Orient. It is a kickoff decision, **never a mid-run interruption**.
- **Do not re-ask for steps this skill mandates.** Invoking plan-exec authorizes you
  to code-review the diff, author the commit message, commit per phase, and hand off
  when context grows. Do them — never stop to ask "should I commit? is the message
  OK? should I review? should I hand off?"
- **Planned migrations and dependency changes are autonomous.** If the plan calls
  for a migration or a dep install/update/downgrade, run it — commit, deps, and
  additive migrations are not gated. A **destructive migration** (data loss) is the
  exception: it stays gated (global DB rule).
- **A mid-run bifurcation that is not destructive → do it and document it** (in the
  plan doc / final summary) so you can review it afterward. Don't stop for a doubt
  that breaks nothing; verify the assumption (investigate, don't guess).
- **Only stop for:** a `deny`-class destructive action (skip + document), an
  un-pre-authorized publish (surface at the end), or a genuine hard blocker you
  cannot resolve (missing credentials, truly unknowable intended behavior).
- **A phase report is not a stop.** Put the status note in the same message as the next
  phase's first tool call; never close a turn by announcing the next step instead of
  taking it (both drew a bare "continue" from the user).
- **On an ambiguous fork you cannot cleanly classify — consult the
  durability-arbiter before stopping.** Launch it with the Agent tool as
  `subagent_type: aidex:durability-arbiter` (definition:
  [`../../agents/durability-arbiter.md`](../../agents/durability-arbiter.md); its own
  `model: sonnet` / `effort: high`, read-only — `model-policy: per-stage`, so the gate's
  depth is pinned rather than inherited from the run it is judging), giving it the
  situation + the run's autonomy surface + the phase's proof (verification output,
  commit SHA). Follow its `CONTINUE` / `ASK` / `STOP` verdict; batch any `ASK` to the
  end. If it errors or returns nothing, apply the rule above and **proceed — never
  block on the arbiter** (it is a forcing function, not a gate).

Otherwise: proceed. The user will redirect if needed.

## Unattended / batch execution (opt-in, gated)

The default path above is **interactive**. For **unattended/batch** runs ("execute the
whole plan while I'm away"), this skill can launch the plan as a durable `Workflow` — each
phase a fresh bounded agent, a two-stage gate per phase, crash-resumable via the journal.
Promote only when the work is **decomposable + machine-verifiable + unattended** and each
phase's real work dwarfs the per-agent floor; the mandatory Orient evaluation is the opt-in.

**Read `${CLAUDE_PLUGIN_ROOT}/skills/plan-exec/references/01-unattended-batch-execution.md`
before promoting anything** — promotion threshold and its measured ~22k/agent cost floor,
the three shipped workflow forms and how to pick one, how to derive `args` from the plan,
the phase tier map, and what happens when a phase fails its gate.

## Workflow

### 0. Orient

1. Read the plan document fully (path is in the user prompt or in
   `.context/plans/`). If multi-file, read `00-index.md` plus the current
   phase file. You may skip **Execution log** entries for already-completed
   phases (canon §Execution log) — they are proof journaling, not spec.
2. Identify: total phases, current phase (first unchecked checkbox), success
   criteria per phase, verification step.
3. **Check whether this plan is bug work.** Plans carry no `type` field; the
   back-link runs the other way, so resolve it by grep:
   `grep -rl "escalated_to: plan/<slug>" .context/backlog/`. If the originating item
   carries `type: bug`, every behavior-changing phase is bound by RED→GREEN
   (`bugfix`): the test is written and fails for the right reason **before** the
   fix, and the GREEN output is that phase's proof. No matching item → carry
   on normally.
4. **Check the prior phase's review evidence.** If a previous phase completed
   this session or an earlier one, confirm its Execution-log entry in
   `00-index.md` carries a `review: <verdict> · <n> findings` line. A missing
   entry means the between-phase code-review was skipped — run it now, on the
   prior phase's diff, before starting the current phase; do not proceed
   silently on an unreviewed phase.
5. **Honor the plan's Isolation surface** if it declares one. If the plan already
   recorded an Isolation note (from `plan`'s Step 5, at plan-creation time), act
   on it directly: run the recorded `worktree.sh new` command before phase 1
   (`--no-infra` only when the plan says code-only); if the project has no worktree
   setup, `EnterWorktree` and note it. If the plan predates this and has no Isolation
   note, run `worktree bootstrap` if `.context/worktrees/00-index.md` does not
   exist yet, then `worktree.sh new` here at Orient, before phase 1. Enter the
   worktree **only if the plan/user authorized it** — do not auto-enter one that was
   not approved. **Before creating any worktree/branch, resolve and state its base branch
   and require explicit confirmation if it is not the repo's default** (worktree's
   branch-base rule) — never fork off the ambient checkout silently.
   No declared surface and no plan-recorded parallelism → run in place.
   **The plan doc stays source-of-truth in the main tree:** a fresh worktree has only
   committed files, so update the plan and record `proof_links` at its main-tree path
   (a gitignored/uncommitted `.context/` plan is absent from the worktree) — see the
   canon's Lifecycle note.
6. **Probe for concurrent work before touching anything.** The user runs
   parallel sessions and worktrees on the same project, and a session blind to
   them will take decisions another session owns. Two commands, seconds:
   `git worktree list` and `git log --all --since="24 hours ago" --oneline`.
   If another live line of work shows — a worktree you did not create, fresh
   commits this session did not make — name it in your first status message,
   keep hands off its files and branches, and route any decision that belongs
   to it back to the user instead of taking it here.
7. **Set the run's spend before phase 1.** Each phase's `tier` says how hard its work
   is; the canon's table
   ([plan-conventions.md](../conventions/references/plan-conventions.md) §"Optional
   phase metadata") maps that to a default model and effort level. Those cells are
   defaults, and this is the one moment to override them — remaining quota, model
   availability, or a cheaper cell the user prefers. State the override and its reason in
   one line and log it to the Execution log; a phase whose cell you changed is not a phase
   whose plan you edited. Ask nothing: an unstated tier is `standard`, and no override is
   the default.
   **A tier row whose cell restricts `tools:` also needs a registered agent definition**, since
   `agent()` has no `tools` option: the toolset travels only as `agentType`, the name of an agent
   file. One definition per tier the plan uses (never per phase) — `model`/`effort` from the canon
   row, `tools:` exactly the row's list, `skills: [aidex:testing]` (a restricted tier has no
   Skill tool), `user-invocable: false` — living in the project's
   `.claude/agents/<name>.md`, or `~/.claude/agents/<name>.md` when several projects share it; those
   are the two locations an agent name resolves from, project first. Pass its name as the batch
   phase's `agentType`. No definition → the phase runs unrestricted and pays ~49k of prefix per call
   instead of ~14k.
   **The definition must be on disk before this session started.** The agent registry is read at
   session start: a stub written mid-session is invisible to the session that wrote it, and a batch
   naming it dies at once with `agent type '<name>' not found`. So check for it here, at Orient —
   if it is missing, write it and hand off; the next session sees it. Two things measured on a real
   batched run (2026-09-20): the restricted implementer's first call read 15,826 `cache_creation`
   against 50,841 for the same run's unrestricted verifier, and the phase's `model` cell won over
   the definition's `model:` line — `agentType` carries the toolset, nothing else.
8. Create a TaskList mirroring the plan's phases so progress is visible.
9. **Front-load the work-list for chained multi-item runs.** A single plan's phases
   are already an ordered queue (walk them). But when this session chains **multiple
   plans/items** (close several plans, then clear backlog), fix the cross-item order
   **once** here — via the `AskUserQuestion` survey → a durable
   `.context/worklists/` work-list (see
   [worklist-conventions.md](../conventions/references/worklist-conventions.md)).
   Then walk it with `worklist-advance.sh` between items instead of pausing to ask
   "what next?". Emergent work (class b) is appended (`--append`) and continued, not
   asked; only a class-(c) fork or the publication gate interrupts.
   **No interactive channel** (`claude -p`, cron): skip the survey, walk the items in
   the order they were given, and record the defaulting in the run's final summary —
   [autonomy-conventions.md § When there is no interactive channel](../conventions/references/autonomy-conventions.md).

### 1. Execute each phase

**Multi-file plan, interactive run:** phases go in waves, not one by one:
[`references/05-launch-waves.md`](references/05-launch-waves.md).

For each phase in order:

1. Implement the tasks in the phase. **Plan code is a sketch, not a paste
   source**: any code block or line reference in the plan was frozen at
   plan-write time — before applying one, read the current file, confirm the
   surrounding code still matches, and check for sibling call-sites/branches
   the plan did not enumerate. The phase's acceptance criteria and gate are
   the contract; the plan's code is illustrative except inside a **Contract**
   block (exact signatures/shapes/DDL), which is binding. A shared-component
   primitive gets one real-framework test per merge-dependent behaviour (attrs/class
   fallthrough, slot defaults), named required in the impl brief.
2. Run the verification step the plan declares (tests, type-check, build,
   manual check). If none is declared, run the minimum that proves the change
   works (relevant test suite + type-check). **Iterate on the selection, not the
   whole suite:** `${CLAUDE_PLUGIN_ROOT}/skills/audit/scripts/affected-tests.sh --command`
   prints one runnable command for the tests covering the phase's diff; exit 3 means
   no selection is available, so run everything and say so. The between-phase
   checkpoint commits on the **selection**, stated rather than silent — say which
   subset ran and that the full suite has not. **The full suite gates the INTEGRATION
   boundary**: the merge, the push, or the end of the run
   (`decision/2026-08-24-full-suite-gate-moves-from-commit-to-integration`). An
   `# INCOMPLETE` selection is the exception — unmapped scope forces the full suite
   in-phase. It also names changed files that **measurably break** and have no E2E —
   write that spec in-phase.
3. If verification fails: fix root cause. After 3 failed attempts on the same
   approach, stop and ask the user.
4. Mark the phase's checkboxes as done in the plan file. **Record the phase's
   proof** — the verification output, the commit SHA, a request/response payload,
   or a screenshot of the flow — in the plan front-matter `proof_links` (or under
   `.context/proofs/<slug>/` for larger captures) per `conventions`
   (`00-global.md` §7.1). Don't mark a phase done you can't show works.
5. **When execution departs from the plan, the plan edit ships in the commit that
   departs** — not in the close-out commit. Otherwise the phase diff is reviewed
   against a stale plan, and the reviewer cannot tell a deliberate change of course
   from an omission. Same coupling `plan-conventions.md` already applies to tests.
   Where the plan is not committable this collapses to editing it before the commit,
   in the same turn: aidex's own repo gitignores `.context/`, and it is the exception
   — 1,135 plan files are tracked across the six fleet repos, 0 here (census
   2026-09-07).
6. **Tests and bugs met mid-phase.** Tests follow `testing`; every delegate brief names
   `aidex:testing`. An in-scope bug is fixed RED-first (`bugfix`). An out-of-scope one —
   found here or reported by a delegate (batch: the result's `outOfScopeBugs`) — is never
   patched inline: route it with its reproduction to a bugfix delegate (`bugfix-opus`
   where defined) or register it ([03](references/03-deferring-emergent-work.md)). If it
   blocks the phase's gate, the phase stays blocked until that fix lands, then resumes.

> **Scoped plans carry a file contract.** When the plan's front-matter says
> `mode: scoped`, its `**Files:**` list is the declared blast radius, written on
> deliberately incomplete investigation — so it will sometimes be wrong. The contract makes
> that **visible, not impossible**: (1) log any file you touch outside the list in the
> Execution log, one line; (2) re-apply the five triage signals (`plan-conventions.md`
> §The five signals) **to that file** — if any flips to `full`, stop and re-triage the
> whole plan with `plan`; (3) independently, if the file list has **doubled**, stop
> and re-triage. Six extra trivial files trip no signal but mean the contract misread the
> change — a failure rule (2) cannot see.

> **Loop (opt-in, per phase only):** if a single phase is mechanical and its verification is a
> pure machine gate (e.g. "make all `<suite>` pass" / "type-check clean"), that one phase may be
> spec'd as a loop via `loop` and run by `/goal`/`ralph-loop` — mirroring the `plan` →
> `loop` pointer at the phase level. **Do not loop the executor itself:** the between-phase
> checkpoint (review/commit/handoff) is judgment work, and irreversible steps
> (push/release/deploy) stay outside any auto-loop and human-gated. (`commit` is
> not irreversible — it is part of the checkpoint, not a gated step.)

### 2. Between-phase checkpoint (MANDATORY)

After each phase passes verification, before starting the next phase, run the shared
checkpoint — **read**
`${CLAUDE_PLUGIN_ROOT}/skills/conventions/references/checkpoint-conventions.md` **and follow its
four moves** (scoped review with its recorded anchor and findings addressed by remedy · commit · register-don't-discuss ·
auto-handoff without asking). It is one canon with two consumers (this skill and the
backlog sweep) and is not restated here; `test_checkpoint_lockstep.sh` fails this file if
it grows its own copy. What is plan-specific:

- **Scope.** `${CLAUDE_PLUGIN_ROOT}/skills/conventions/scripts/resolve-review-scope.sh --files working-diff`,
  or `--base <phase-start-sha> branch-vs-main` for a phase that spans commits — and
  `${CLAUDE_PLUGIN_ROOT}/skills/conventions/references/review-scope-conventions.md` owns which
  reviewer covers which scope. **Exit 3 is an empty scope, never a passing review.**
- **Where the evidence goes.** The Execution log in the plan's `00-index.md` takes the
  `review: <verdict> · <n> findings · scope=<scope> anchor=<anchor>` line before the commit.
- **A UI phase does not close on prose.** Before the commit, run
  `${CLAUDE_PLUGIN_ROOT}/skills/plan-exec/scripts/check-ui-evidence.sh <plan file> <phase>`
  (the plan or its `00-index.md`). It owns which phase is UI and the `ui-*` lines
  "verified" needs; non-UI exits 0. **Exit 1 or 2 keeps the phase open.**
  An unattended run may queue a non-final phase's page as `pending-owner` (script header); the last phase fails while any is open.
- **Deferrals** use `register-item.sh --origin plan --plan <this plan>`
  ([`references/03-deferring-emergent-work.md`](references/03-deferring-emergent-work.md)).
- **Which model runs which step** — orchestrate, implement, and do the mechanical work
  with different models
  ([`references/04-model-tiering.md`](references/04-model-tiering.md)).
- **The seed's `slug:` line is the PLAN's name** plus the phase — the name existed before
  the first handoff and does not move. `CHARTER` comes from the plan's name and goal.

### 3. Final phase

After the last phase:

1. Run the full verification the plan declares (or the project's standard
   pre-deploy check: tests + build + lint).
2. Code-review and commit as above.
3. If the plan implies a release (user-facing changes, feature complete):
   surface the project's release command as an option (detect it — many
   projects expose a `/release`-style command). Do not run it without explicit
   user approval — releases are deploy-coupled.
4. Update the plan document: mark all phases complete, add a closing note
   with the final commit SHAs if useful.
5. **Close out the run**: tear down isolation if a worktree was entered at Orient, log the
   worktree usage line, suggest a coverage sweep if the plan touched mapped src paths,
   run guided human verification — it emits a proof artifact, and a plan with
   nothing human-visible skips it by recording the reason, never by omission —
   reconcile deferrals to a `BL-NNN` or a `CLOSE`, and fire the notifier. **Read**
   `${CLAUDE_PLUGIN_ROOT}/skills/plan-exec/references/02-close-out.md`
   **and follow it step by step** — each step has a guard and an ordering that
   matter, and doing them from memory is how a worktree survives its plan.

## Per-project adjustments

This skill ships stack-agnostic defaults; detect the project's own conventions:

- **Review/commit/release commands.** Use the project's own slash commands or
  helpers (`.claude/`, CLAUDE.md); never assume a name.
- **Stricter project rules.** Read the project's CLAUDE.md (and memory) before the
  first phase — test runners, commit style, version bumps, release gates.

If the project's CLAUDE.md or memory contradicts this skill, the project wins.

## What this skill does NOT do

- It does not create plans (use `plan`).
- It does not skip verification to move faster — every phase is verified.
- It does not deploy or release without explicit user approval.
- It does not run E2E tests against dev environments — use the project's
  isolated test runner if E2E is required by a phase.

Files in this skill

  • SKILL.md17.5 KB
  • assets/workflows/fan-out-with-gate.workflow.js17.6 KB
  • assets/workflows/pipeline-with-gate.workflow.js14.8 KB
  • assets/workflows/review-with-gate.workflow.js16.5 KB
  • evals/native/execute-phase-one/case.yaml124 B
  • evals/native/execute-phase-one/graders/plan-exec-conventions.md1.1 KB
  • evals/native/execute-phase-one/graders/skill-fired.md74 B
  • evals/native/execute-phase-one/graders/slugify-test-file-created.md56 B
  • evals/native/execute-phase-one/prompt.md366 B
  • evals/native/execute-phase-one/setup.sh3.6 KB
  • evals/trigger_eval.json5.5 KB
  • references/01-unattended-batch-execution.md13 KB
  • references/02-close-out.md5.7 KB
  • references/03-deferring-emergent-work.md2.3 KB
  • references/04-model-tiering.md3 KB
  • tests/test-agents-set-model.sh1 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…