Skip to content
Back to skills

Autopilot

ASecurity

Use this skill when running an autonomous session-orchestration loop. Chains session-start → session-plan → wave-executor → session-end for N iterations with all 10 kill-switches (SPIRAL, FAILED wave, carryover > 50%, max-hours, max-sessions, resource-overload, token-budget, stall-timeout, sub-threshold confidence, user-abort). Reads Mode-Selector output (Phase B) to decide auto-execute vs. fallback. Writes one autopilot.jsonl record per loop run. Phase C scaffold (issue #277); implementation...

  • 53 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added May 28, 2026
ai-agentsgogitapi

Works with

  • claude code
  • cli
  • api

Security analysis

A100/100

Scanned October 4, 2026

npx -y skills add Kanevry/session-orchestrator --skill autopilot --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Autopilot?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Autopilot
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/kanevry-autopilot/badge)](https://www.skillsdirectory.com/skills/kanevry-autopilot)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: autopilot
description: >
  Use this skill when running an autonomous session-orchestration loop. Chains session-start → session-plan →
  wave-executor → session-end for N iterations with all 10 kill-switches (SPIRAL, FAILED
  wave, carryover > 50%, max-hours, max-sessions, resource-overload, token-budget,
  stall-timeout, sub-threshold confidence, user-abort). Reads Mode-Selector
  output (Phase B) to decide auto-execute vs. fallback. Writes one autopilot.jsonl
  record per loop run. Phase C scaffold (issue #277); implementation lives in
  scripts/lib/autopilot.mjs (Phase C-1 follow-up).
user-invocable: true
argument-hint: "[--headless] [--verbose] [--max-sessions=N] [--max-hours=H] [--confidence-threshold=0.X] [--dry-run]"
tags: [phase-c, autopilot, autonomous, loop]
model: sonnet
---

# Autopilot Skill

## Invocation

The user invokes `/autopilot` with arguments: **$ARGUMENTS**. This is autonomous
session-orchestration mode — a top-level invocation only, never from inside a running
session.

Parse `$ARGUMENTS` before anything else. Unrecognized flags are ignored; out-of-range
values silently clamp to bounds. Use `parseFlags` from `scripts/lib/autopilot.mjs` for
canonical parsing — never re-implement clamping inline. The loop flags
(`--max-sessions`, `--max-hours`, `--confidence-threshold`, `--dry-run`) and their
defaults/bounds are tabled once in § Command Surface below. Two further flags belong to
the invocation surface only:

| Flag | Default | Meaning |
|------|---------|---------|
| `--headless` | `false` | Run via the standalone driver `scripts/autopilot.mjs`, which spawns `claude -p '/session <mode>'` as a child process per iteration. Required for unattended walk-away mode. Without it, `/autopilot` runs the in-process driver inside the current chat session (see § Production Wiring). |
| `--verbose` | `false` | Verbose driver output. |

### Headless (`claude -p`)

Headless requirements:

- Anthropic API key configured for `claude -p` (Claude Code CLI auth).
- `scripts/autopilot.mjs` starts each child with `--session-id <fresh uuid>` and, after it
  exits, reads the one canonical `sessions.jsonl` record whose `raw_session_id` is that uuid
  to construct the `sessionRunner` return shape — the spawned session must complete cleanly
  and append its record (session-end Phase 3.7 handles this). No such record, or more than
  one, stops the loop with `failed-wave` instead of reading the ledger tail (#1457). So does
  a unique record that is only a backfill stub (no waves/agents/token totals — the child never
  closed): its token usage is unknown, not 0, so it is never counted as a healthy 0-token
  iteration (#1457 F2).
- `AUTOPILOT_RUN_ID` env var is set in the child's env, but **nothing reads it yet**
  (measured 2026-10-01: `rg -n "AUTOPILOT_RUN_ID" scripts hooks skills` finds only the writer
  in `scripts/autopilot.mjs` and this line). CLI-driven iterations therefore carry no
  `autopilot_run_id` on their `sessions.jsonl` record yet; join them to `autopilot.jsonl` via
  the `sessions` list there, not via that field.

Do NOT re-implement loop logic inline — this skill and `scripts/lib/autopilot.mjs` are
authoritative. Kill-switches are enforced by `scripts/lib/autopilot.mjs`, not inline by
Claude. The runtime writes ONE record to `.orchestrator/metrics/autopilot.jsonl` per
invocation via atomic tmp+rename; no other code path appends to that file.

## Phase 0.5: Parallel-Aware Preamble

> Skip silently when `persistence: false` in Session Config.

Before any Phase 1 work, run the parallel-aware preamble per `skills/_shared/parallel-aware-preamble.md`. The preamble detects other active sessions in the worktree-family via `findPeers(repoRoot, { mySessionId })`, classifies the caller's mode via `classifyMode(callerMode)` against the exclusivity-matrix, and either:

- Returns `PASS_THROUGH` (no other session / `always-ok` mode) → continue to Phase 1
- Returns `EXCLUSIVE_BLOCKED` → fires Exclusive-Conflict AUQ from `skills/_shared/parallel-aware-auq.md`
- Returns `PROMOTION_OFFER` → fires Worktree-Promotion AUQ (via `enterWorktree()` from `scripts/lib/autopilot/worktree-pipeline.mjs` — see `parallel-aware-auq.md` outcome-handling)

On any non-PASS_THROUGH outcome that does not result in immediate exit, append a Deviation to STATE.md via `appendDeviationOnDisk(repoRoot, isoTimestamp, message)` from `scripts/lib/state-md.mjs`.

**Implementation reference:** `skills/_shared/parallel-aware-preamble.md § Implementation`.
**AUQ reference:** `skills/_shared/parallel-aware-auq.md`.

## Status

**Phase C-1.b complete (2026-04-25, issues #295 + #300).** Runtime at
`scripts/lib/autopilot/kill-switches.mjs:18-32` (the frozen `KILL_SWITCHES` enum is
SSOT) enforces all 10 kill-switches:

- **Pre-iteration (6, #295 + #355):** `max-sessions-reached`, `max-hours-exceeded`,
  `resource-overload`, `low-confidence-fallback` (with iter-1-fallback /
  iter-2+-exit asymmetry), `user-abort`, `token-budget-exceeded` (cumulative tokens
  ≥ `--max-tokens`).
- **Post-iteration (1, ADR-364 §3):** `stall-timeout` — no progress marker in
  `autopilot.jsonl` within the threshold (default 600s; missing file → no kill).
- **Post-session (3, #300):** `spiral`, `failed-wave`, `carryover-too-high`.
  Read schema-canonical fields off the `sessionRunner` return shape:
  `agent_summary.{spiral, failed}` (numeric counts) and `effectiveness.{carryover,
  planned_issues}`. Absent fields → no kill (forward-compatible: a `sessionRunner`
  that does not yet emit those fields silently no-ops the post-session gates).

Atomic `autopilot.jsonl` writer (tmp+rename, schema_version 1) and silent-clamp
`parseFlags` shipped in C-1. `autopilot_run_id` is passed into `sessionRunner`
via `args.autopilotRunId`; production callers MUST persist it into the per-iteration
`sessions.jsonl` record (additive optional field, schema_version 1 compatible).
See `skills/wave-executor/SKILL.md § Return Shape Contract` and
`skills/session-end/SKILL.md § Phase 3.7`.

## Purpose

Autopilot collapses the per-session attention cost when Mode-Selector is confident enough
to make routine decisions autonomously. A productive day commonly ships 3–7 sessions; each
manual session-start costs the user 10–60 seconds of context-switch attention. When the
session is genuinely routine (mechanical refactor, post-merge housekeeping, repeated
follow-ups from a planned epic), that attention cost is pure overhead.

`/autopilot` reads the Mode-Selector recommendation, executes the recommended session if
confidence clears the threshold, then loops — checking kill-switches between iterations.
The user invokes the loop once and walks away; autopilot stops itself when work runs out
or quality degrades.

This is **opt-in by design**: autopilot never starts itself. The user must run
`/autopilot` explicitly. Configuration thresholds (`--max-sessions`, `--max-hours`,
`--confidence-threshold`) are CLI flags, not Session Config defaults — the user signals
intent for THIS run, not a standing policy.

## Command Surface

```
/autopilot [--max-sessions=N] [--max-hours=H] [--confidence-threshold=0.X] [--max-tokens=N] [--dry-run]
```

| Flag | Default | Bounds | Meaning |
|------|---------|--------|---------|
| `--max-sessions` | `5` | 1..50 | Iteration cap (graceful exit when reached) |
| `--max-hours` | `4.0` | 0.5..24.0 | Wall-clock budget for entire loop |
| `--confidence-threshold` | `0.85` | 0.0..1.0 | Minimum `selectMode` confidence for auto-execute |
| `--max-tokens` | `0` (off) | 0..10000000 | Token budget for `token-budget-exceeded`. Unit: cumulative **subagent OUTPUT** tokens, summed from each `sessions.jsonl` record's `total_token_output` (falling back to `total_tokens`, which is subagent-only as well) — coordinator tokens are not counted. Off by default until calibrated: one measured deep session's subagents produced 450 499 output tokens, so a six-figure budget stops the loop after about one deep session. An iteration whose record carries no token figure adds nothing and is counted as `tokens_unknown_sessions` (state, return value, `autopilot.jsonl`, and the run summary line on stdout, which prints `tokens=<sum>` and — when non-zero — `tokens_unknown_sessions=<n>`), so with any such iteration the sum is a LOWER bound and the budget can be overshot |
| `--dry-run` | `false` | — | Print planned iterations without executing |

Out-of-range values silently clamp to bounds. `--dry-run` exits after printing — never
invokes session lifecycle.

## Loop Semantics

```
state := { iterations_completed: 0, started_at: now(), kill_switch: null, sessions: [] }

WHILE state.iterations_completed < max-sessions:
  # Pre-iteration kill-switches (6)
  IF aborted: kill_switch := 'user-abort'; break
  IF state.iterations_completed >= max-sessions:
    kill_switch := 'max-sessions-reached'; break
  IF (now() - state.started_at) > max-hours:
    kill_switch := 'max-hours-exceeded'; break
  IF max-tokens > 0 AND cumulative_tokens_used >= max-tokens:
    kill_switch := 'token-budget-exceeded'; break
  IF resource_verdict() == 'critical' AND peer_count() > autopilot-peer-abort:
    kill_switch := 'resource-overload'; break

  recommendation := mode-selector.selectMode(<live signals from session-start Phase 7.5>)

  IF recommendation.confidence < confidence-threshold:
    IF state.iterations_completed == 0:
      fallback_to_manual()  # iteration 1: hand off cleanly to manual /session flow
    ELSE:
      kill_switch := 'low-confidence-fallback'  # iteration 2+: exit, let user decide
    break

  cap := resource_adaptive_cap()
  session_result := run_session(mode=recommendation.mode, agents_per_wave_cap=cap)
  state.sessions.append(session_result.session_id)

  # Post-iteration kill-switch (1)
  IF stalled(autopilot.jsonl) >= stall-timeout: kill_switch := 'stall-timeout'; break
  # Post-session kill-switches (3)
  IF session_result.spiral_detected: kill_switch := 'spiral'; break
  IF session_result.failed_waves > 0: kill_switch := 'failed-wave'; break
  IF session_result.carryover_ratio > 0.50: kill_switch := 'carryover-too-high'; break

  state.iterations_completed += 1

write_autopilot_jsonl(state, kill_switch)
print_summary(state, kill_switch)
```

**Atomicity rule:** iteration boundaries are atomic. A session must complete (`/close`
including the post-session writes) before the next iteration starts. Autopilot does NOT
abort sessions mid-flight; kill-switches are checked AFTER each session completes.

## Kill-Switches

All 10 kill-switches, grouped by check phase (mirrors the `KILL_SWITCHES` enum in
`scripts/lib/autopilot/kill-switches.mjs:18-32`):

| Kill-switch | Phase | Trigger | Recovery hint |
|-------------|-------|---------|---------------|
| `max-sessions-reached` | pre-iteration | `iterations_completed >= --max-sessions` | Graceful — not an error. |
| `max-hours-exceeded` | pre-iteration | Wall-clock exceeds `--max-hours` | Re-run with higher `--max-hours` or address slow waves. |
| `resource-overload` | pre-iteration | `verdict==critical AND peers > autopilot-peer-abort` | Wait for peer sessions to complete or close them. |
| `low-confidence-fallback` | pre-iteration | `confidence < threshold` (iteration 2+) | Re-run with lower `--confidence-threshold` or run next session manually. |
| `user-abort` | pre-iteration | Ctrl+C / Esc (AbortSignal) | Re-run when ready. |
| `token-budget-exceeded` | pre-iteration | `cumulative_tokens >= --max-tokens` (#355) | Re-run with a higher `--max-tokens` budget or split the work. |
| `stall-timeout` | post-iteration | No progress marker in `autopilot.jsonl` within threshold (ADR-364 §3; default 600s) | Inspect the stalled iteration; missing telemetry file is NOT a kill. |
| `spiral` | post-session | wave-executor spiral detection fires (`agent_summary.spiral > 0`) | Triage the spiraling wave manually; autopilot will not retry. |
| `failed-wave` | post-session | Any wave reports `agent_summary.failed > 0` | Investigate failure mode (test contract drift, env issue). Re-run after fix. |
| `carryover-too-high` | post-session | `carryover/planned > 0.50` | Last session under-delivered. Reduce scope or split issues before resuming. |

## Resource-Adaptive Concurrency

Autopilot does NOT hard-block on peer Claude processes. It adapts `agents-per-wave` cap
per iteration based on the most-restrictive resource signal.

| Tier | RAM free | Swap | Peers | macOS memory_pressure | cap |
|------|----------|------|-------|------------------------|-----|
| green | ≥ 6 GB | < 1 GB | ≤ 2 | ≥ 30% free | Session Config default |
| warn | 4–6 GB | 1–2 GB | 3–4 | 15–30% free | 4 |
| degraded | 2–4 GB | 2–3 GB | 5–6 | 5–15% free | 2 |
| critical | < 2 GB | > 3 GB | > 6 | < 5% free | 0 (coord-direct) |

**Most-restrictive-signal-wins:** `[ram=8GB, swap=0, peers=7]` → critical (peer rule wins).

Defaults are conservative initial estimates. Phase C-3 follow-up calibrates the swap and
memory_pressure thresholds against real autopilot-run effectiveness data.

## Production Wiring

Phase C-1 ships `runLoop` as a pure controller. Phase C-1.c ships `buildLiveSignals` as
the canonical signals-assembly helper. This section documents the **in-process driver
protocol** (Option B from #301): how Claude — running as the coordinator in a chat
session — drives `runLoop` between manual `/session` invocations. The headless wrapper
(Option A, `scripts/autopilot.mjs` CLI spawning `claude -p`) is selected with `--headless`
— see § Invocation.

### Dependency-Injection Contract

`runLoop` requires four injected dependencies:

| Field | Signature | Source |
|---|---|---|
| `modeSelector` | `() => Promise<{mode, confidence, rationale?}>` | wraps `selectMode(await buildLiveSignals())` |
| `sessionRunner` | `({mode, autopilotRunId}) => Promise<{session_id, agent_summary?, effectiveness?}>` | wraps a `/session <mode>` invocation; reads the child's own record (matched on `raw_session_id`) to construct return value |
| `resourceEvaluator` | `() => {verdict}` | calls `evaluate(cachedProbeSnapshot, thresholds)` from `resource-probe.mjs` over a snapshot `peerCounter` refreshed on the prior iteration — never calls `probe()` itself, which is what keeps it synchronous |
| `peerCounter` | `() => Promise<number>` | returns `peers.length` from `detectPeers({ sessionId, freshnessMin: 15 })` (a SESSION count, not a process count — see `host-resources.md` HR-103) while refreshing the cached `probe()` snapshot `resourceEvaluator` reads |

`abortSignal` is optional (Ctrl+C / Esc → `user-abort` kill-switch).

### In-Process Driver Skeleton

```js
import { runLoop, parseFlags } from '$PLUGIN_ROOT/scripts/lib/autopilot.mjs';
import { buildLiveSignals } from '$PLUGIN_ROOT/scripts/lib/build-live-signals.mjs';
import { selectMode } from '$PLUGIN_ROOT/scripts/lib/mode-selector.mjs';
import { probe, evaluate } from '$PLUGIN_ROOT/scripts/lib/resource-probe.mjs';
import { detectPeers } from '$PLUGIN_ROOT/scripts/lib/session-registry.mjs';

const flags = parseFlags(process.argv.slice(2));

const modeSelector = async () => {
  // Each iteration rebuilds signals from current disk state. STATE.md will be
  // freshly idle-reset by the previous /close, sessions.jsonl will have the
  // new tail entry, etc. This is the contract: live signals every iteration.
  // `repoRoot` is passed explicitly (#1071) rather than left to the helper's
  // cwd default. Do NOT hand-write a `backlogLimit` here — the window default
  // lives once, in backlog-scan.mjs (`DEFAULT_BACKLOG_LIMIT`); see
  // skills/session-start/phase-7-5-mode-selector.md for the same contract.
  const signals = await buildLiveSignals({ repoRoot: process.cwd() });
  return selectMode(signals);
};

let cachedProbeSnapshot = null;   // written by peerCounter, read by resourceEvaluator

const resourceEvaluator = () => {
  // Synchronous by contract: never calls probe() itself (it is async) — reads
  // the snapshot peerCounter refreshed on the prior iteration.
  if (cachedProbeSnapshot === null) {
    return { verdict: 'warn', reasons: ['probe not yet available'], recommended_agents_per_wave_cap: null };
  }
  return evaluate(cachedProbeSnapshot, thresholds);
};

const peerCounter = async () => {
  // A SESSION count, not a process count (host-resources.md HR-103) —
  // claude_processes_count runs ~6 processes per session. `ownSessionId` is
  // this coordinator's registry id, so detectPeers() excludes it.
  const [peers, snapshot] = await Promise.all([
    detectPeers({ sessionId: ownSessionId, freshnessMin: 15 }),
    probe(),
  ]);
  cachedProbeSnapshot = snapshot;
  return peers.length;
};

const sessionRunner = async ({ mode, autopilotRunId }) => {
  // The coordinator (Claude) invokes /session <mode> manually here. After the
  // session completes (/close runs, sessions.jsonl appended), this function
  // reads the tail entry and projects it into the runLoop return-shape.
  // CAUTION: a tail read has the peer-append error class — a parallel session
  // that appends after this one supplies ITS record (and token usage). A real
  // run should read its own record by raw session id, as scripts/autopilot.mjs
  // does (readOwnSession → raw_session_id match, #1457).
  const tail = readSessionsJsonlTail(1);   // last line, normalized
  return {
    session_id: tail.session_id,
    agent_summary: tail.agent_summary,    // {complete, partial, failed, spiral}
    effectiveness: tail.effectiveness,    // {planned_issues, carryover, completion_rate, ...}
  };
};

const result = await runLoop({
  ...flags,
  modeSelector,
  sessionRunner,
  resourceEvaluator,
  peerCounter,
});
```

### Why In-Process First

The in-process driver has Claude (the coordinator) call `/session <mode>` between
`runLoop` iterations, with `runLoop` orchestrating the kill-switches. Trade-offs:

- **Pro:** zero new infra. Reuses canonical kill-switch logic. Validates `buildLiveSignals`
  against real Phase 7.5 swap before headless complexity. Each iteration carries
  inter-session memory through STATE.md / sessions.jsonl / learnings.
- **Con:** not truly autonomous — Claude must stay in the chat. The in-process
  driver does not deliver walk-away UX; that is the shipped headless driver
  (`--headless`, `scripts/autopilot.mjs`, Phase C-5; see § Headless Driver Wiring).

### Headless Driver Wiring (Option A — `scripts/autopilot.mjs`)

The standalone headless driver invoked via `--headless` (see § Invocation) wires all
four required `runLoop` dependencies (`modeSelector`, `sessionRunner`,
`resourceEvaluator`, `peerCounter`) plus the optional `abortSignal` to production
sources — distinct from, and more concrete than, the in-process skeleton above:

- `sessionRunner` — spawns `claude -p '/session <mode>' --session-id <uuid>` as a child
  process; after it exits cleanly, reads the one canonical `sessions.jsonl` record whose
  `raw_session_id` is that uuid to construct the return shape
  `{session_id, agent_summary?, effectiveness?, usage?}` (`scripts/autopilot.mjs` `sessionRunner`).
- `resourceEvaluator` — calls `evaluate(cachedProbeSnapshot, thresholds)` from
  `scripts/lib/resource-probe.mjs`, reading a snapshot refreshed by `peerCounter` on the
  prior iteration so the function itself stays synchronous, per the `runLoop` contract
  (`scripts/autopilot.mjs` `resourceEvaluator`).
- `peerCounter` — calls `detectPeers({ sessionId: autopilotRunId, freshnessMin: 15 })` from
  `scripts/lib/session-registry.mjs` AND refreshes the cached `probe()` snapshot in the
  same call, returning `peers.length` (`scripts/autopilot.mjs` `makePeerCounter`).
- `abortSignal` — an `AbortController` aborted on the process's `SIGINT` handler.

### `autopilot_run_id` Propagation

When `runLoop` invokes `sessionRunner({mode, autopilotRunId})`, the per-iteration
`sessions.jsonl` record MUST carry `autopilot_run_id: <id>`. session-end Phase 3.7
writes this field. Manual sessions write `null` or omit it — readers treat both
identically per the v1 schema additive convention. See
`skills/session-end/session-metrics-write.md`.

### Acceptance Signals

A live `/autopilot` invocation against this wiring produces a non-zero `confidence`
recommendation when at least one signal source is populated (state-md rec fields,
sessions.jsonl tail, learnings, or backlog). Confidence at 0.0 with all four sources
populated is a Mode-Selector heuristic bug (file as `[Mode-Selector v1.x quirk]` issue),
not an autopilot bug.

## Pre-Loop Verdict Gate (dispatcher → autopilot handoff — #682)

When the cross-repo dispatcher (`skills/dispatcher/SKILL.md`) routes into an autopilot
launch, a **pre-loop suitability verdict** decides whether the launch may proceed WITHOUT
per-selection operator confirmation. This is distinct from — and runs BEFORE — the loop's
10 kill-switches:

- **The verdict gate is a PRE-LAUNCH decision.** It is computed ONCE, at the
  dispatcher → `runLoop` handoff, before the first iteration starts. It answers "may I
  launch this repo autonomously, or must I ask first?" — NOT "should I stop the running
  loop?"
- **The 10 kill-switches are PER-ITERATION and reused UNCHANGED.** Once `runLoop` starts,
  the frozen `KILL_SWITCHES` enum (`scripts/lib/autopilot/kill-switches.mjs:18-32`)
  governs when the loop stops, exactly as documented above. The verdict gate adds NO new
  kill-switch, modifies NONE of the existing 10, and does not re-implement any of them.
- **The gate engine is `computeSuitabilityVerdict(deps)`** from
  `scripts/lib/autonomy/suitability.mjs` — a pure four-gate AND (confidence ≥ floor;
  kill-switch fired-rate < 0.2 over the recent runs, omitted below 5 runs; CI ≠ red;
  resource ≠ critical). The dispatcher gathers every signal and passes it in (DI); the
  engine reads no files.
- **Kill-switch history feeds G2 via `readRecentAutopilotRuns`** from
  `scripts/lib/autopilot/recent-runs.mjs`, which reads THIS repo's
  `.orchestrator/metrics/autopilot.jsonl` (newest-last, never throws). The verdict's G2
  gate counts those records and reads each one's persisted `kill_switch` field — it never
  re-enumerates or re-derives the switches; it reads the history the loop already wrote.
- **FAIL-CLOSED launch wiring:** the dispatcher launches without confirmation ONLY when
  `autonomy === 'autonomous-gated'` AND `verdict.suitable === true`. Every other case
  (any non-`autonomous-gated` dial, a CI-red / resource-critical / low-confidence verdict)
  informs the operator and asks before launch. `resolveDispatcherAutonomy` defaults to
  `'off'` when unconfigured, so an absent config forces inform + ask. See
  `skills/dispatcher/SKILL.md § Phase 1.5` for the full sourcing table and invariant.
- **`null` signals are honest, not failures (NICE-b).** On a CI-fetch or resource-probe
  failure the dispatcher passes `ci = null` / `resourceVerdict = null` (not a synthesized
  `{ status: undefined }` or a fabricated `'green'`). Each `null` ⇒ the gate passes + warns
  — it surfaces a warning the operator sees, it does not block on its own.
- **forcedFail is reachable end-to-end (NICE-c).** When CI is red OR resource is critical,
  `verdict.suitable === false` REGARDLESS of confidence (the engine words the rationale
  `FORCED: CI red` / `FORCED: resource critical`), so even under `autonomous-gated` the launch
  falls to inform + ask. Reachability depends on the dispatcher wiring the live signals
  through (CI wrapped as `{ status }`, the real resource verdict string) rather than masking
  them — see `skills/dispatcher/SKILL.md § Phase 1.5`.

This gate does NOT change the loop. Autopilot remains opt-in, and the kill-switch contract
is unchanged — the verdict gate only governs HOW the loop is entered (auto vs. confirm).

## Telemetry

One record per `/autopilot` invocation, written to `.orchestrator/metrics/autopilot.jsonl`
via atomic tmp + rename. See "Autopilot Loop" (#277; archived in the private Meta-Vault) § Output for the
full schema.

Each iteration's `sessions.jsonl` entry gets an additional optional field
`autopilot_run_id` (string or null) so retros can join across the two files without
schema changes.

Manual sessions write `autopilot_run_id: null` (or omit the field — both treated
identically by readers per the v1 schema additive convention).

## Integration with Other Skills

- **`mode-selector.mjs::selectMode`** — sole source of mode + confidence per iteration.
  Autopilot does not implement its own mode logic; v1.x quirks affect autopilot exactly
  as they affect manual session-start.
- **`resource-probe.mjs::probe + evaluate`** — extended in Phase C-2 with swap and
  memory_pressure signals. Existing consumers (manual session-start, wave-executor)
  benefit from the new signals automatically.
- **`session-start` / `session-plan` / `wave-executor` / `session-end`** — invoked
  unmodified. Autopilot is a controller around the existing session lifecycle, not a
  replacement.
- **`session-registry.mjs`** — peer-count signal source. Autopilot reads but does not
  write to the registry beyond the standard hook.
- **`mode-selector-accuracy`** — autopilot iterations write accuracy learnings exactly
  like manual sessions (Phase B-4 contract). The `chosen` field reflects autopilot's
  auto-execute decision, which equals `recommendation.mode` when confidence ≥ threshold.

## Critical Rules

- **Never auto-merge or auto-push beyond `/close` defaults.** `/close` already pushes to
  origin; autopilot does not add PR creation, merge, or force-push behavior.
- **Iteration boundaries are atomic.** Never abort a running session to start a new one.
- **Kill-switches checked AFTER each session.** Even if a kill-switch will fire after
  iteration N, iteration N completes cleanly first.
- **Iteration 1 sub-threshold falls back to manual; iteration 2+ exits with kill-switch.**
  This asymmetry is intentional — see PRD Q8.
- **`autopilot.jsonl` is the SOLE writer's responsibility of `autopilot.mjs`.** Other
  skills must not append to or rewrite this file.
- **Mode-Selector contract is read-only here.** Autopilot does not modify `selectMode`
  output, does not re-rank alternatives, does not patch confidence values.

## Anti-Patterns

- Do not invoke `/autopilot` from inside a running session. The skill is a top-level
  command; nested invocation is undefined behavior.
- Do not modify `autopilot.jsonl` schema additively without bumping `schema_version`.
  Readers MUST treat unknown fields as a forward-compat signal, not as corruption.
- Do not bypass `selectMode` to force-run a specific mode. If you want to run a specific
  mode, use `/session [mode]` manually — that is autopilot's fallback path.
- Do not lower `--confidence-threshold` below 0.5 in production. The Mode-Selector
  fallback table treats `< 0.5` as suggestion-only; autopilot at that threshold becomes
  a random-walk over modes.
- Do not implement kill-switch logic in this skill. Kill-switch enforcement is in
  `scripts/lib/autopilot.mjs`. The skill documents the contract; the runtime enforces it.

## Configuration

Single-story `/autopilot` takes no Session Config block. Multi-story
(`autopilot.bg-isolation`, `--multi-story`, `--deconflict-paths`) was removed in
4.0.0 — see `docs/migration-v4.md` and `docs/session-config-reference.md` §
"Autopilot Multi-Story (#431) — removed".

## References

- PRD: "Autopilot Loop" (#277; archived in the private Meta-Vault)
- Implementation (Phase C-1 + C-1.b): `scripts/lib/autopilot.mjs` — exports `runLoop`, `parseFlags`, `writeAutopilotJsonl`, `KILL_SWITCHES`, `FLAG_BOUNDS`, `SCHEMA_VERSION`, `DEFAULT_PEER_ABORT_THRESHOLD`, `DEFAULT_JSONL_PATH`, `DEFAULT_CARRYOVER_THRESHOLD`
- Tests (Phase C-1 + C-1.b): `tests/lib/autopilot.test.mjs`
- Mode-Selector contract: `skills/mode-selector/SKILL.md`
- Resource probe: `scripts/lib/resource-probe.mjs`
- Session registry: `scripts/lib/session-registry.mjs`
- Wave-executor return shape: `skills/wave-executor/SKILL.md § Return Shape Contract`
- Sessions.jsonl writer: `skills/session-end/session-metrics-write.md`
- Pre-loop verdict gate (#682): `scripts/lib/autonomy/suitability.mjs` (`computeSuitabilityVerdict`) · `scripts/lib/config/dispatcher-autonomy.mjs` (`resolveDispatcherAutonomy`) · `scripts/lib/autopilot/recent-runs.mjs` (`readRecentAutopilotRuns`) · `skills/dispatcher/SKILL.md § Phase 1.5`
- Epic: [#271 v3.2 Autopilot — Autonomous Session Orchestration](https://github.com/Kanevry/session-orchestrator/issues/271)
- Issues: [#277 Phase C scaffold](https://github.com/Kanevry/session-orchestrator/issues/277), [#295 Phase C-1 runtime](https://github.com/Kanevry/session-orchestrator/issues/295), [#300 Phase C-1.b follow-up](https://github.com/Kanevry/session-orchestrator/issues/300)
- Phase A PRD: "STATE.md Recommendations Contract" (#271; archived in the private Meta-Vault)
- Phase B PRD: "Mode Selector" (#276; archived in the private Meta-Vault)

## Open Questions (Phase C-1 to resolve)

- `--confidence-threshold=auto` — let autopilot self-tune from accumulated
  `mode-selector-accuracy` learnings? Requires ≥ 20 accuracy learnings before useful.
- STATE.md `autopilot-active: true` field — should other Claude sessions detect via the
  session-registry and refuse to start during an autopilot run? Dogfooding will inform.
- `failed-wave` granularity — distinguish "agent failed but was retried successfully"
  from "wave ended with un-recovered failures"? Requires wave-executor schema audit.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…