Skip to content
Back to skills

Autonomous Work

ASecurity

Use for broad autonomous work or recovering a prior project mission on "continue", "resume", "pick up", "продолжи работу". Recover scope and original user grants before selecting work. Bounded tasks stay bounded; continuation does not create outward, destructive, or spending authority.

  • 3 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 22, 2026
ai-agentspythongoshellbashsecuritydocumentation

Works with

  • terminal
  • cli

Security analysis

A100/100

Pro scans all 2 files and shows the line behind each finding

Scanned October 6, 2026

npx -y skills add oleg494/coding-kit --skill autonomous-work --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Autonomous Work?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Autonomous Work
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/oleg494-autonomous-work/badge)](https://www.skillsdirectory.com/skills/oleg494-autonomous-work)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: autonomous-work
description: 'Use for broad autonomous work or recovering a prior project mission on "continue", "resume", "pick up", "продолжи работу". Recover scope and original user grants before selecting work. Bounded tasks stay bounded; continuation does not create outward, destructive, or spending authority.'
license: MIT
metadata:
  version: "4.7.0"
---

# Autonomous Work

Broad authorization to choose and continue useful work — executed without
manufacturing a task for the user and without pretending completion.

## Activation & scope

- **Activates on broad authorization**: "do useful work", "pick what matters",
  "keep going without asking", "работай сам", "делай что полезно". The user has
  authorized *selection* and *continuation*, not a specific task list.
- **Bounded requests are unchanged**: "fix this bug", "review this PR",
  "add this flag" keep their existing scope. Autonomous work adds no scope to a
  request that already names its scope.
- **Opt-in, not a mode override**: this skill is not always-on, and it does not
  replace `STRICT_AUDIT` or `EXPLORATORY_PROTOTYPE`. Under a read-only or
  strict-audit instruction, autonomy selects only read-only/audit work; it
  never upgrades permissions, and it never treats autonomy as a reason to
  weaken a constraint the user or host set.
- If no useful in-scope work remains, say so plainly and stop. Idle is an
  honest outcome; invented work is not.

### Activation and continuation

Before selecting or executing a new autonomous mission, establish one explicit boundary agreement with the user. This happens on autonomous activation, not session startup, bounded requests, or every iteration.

Continuation language (`continue`, `resume`, `pick up`, `продолжи работу`) is a recovery request when a prior project mission may exist. Recover first; do not treat it as a new mission or a symptom-only task.
If current context already establishes the task and authority, reuse it;
do not rediscover the mission for a self-contained bounded continuation.

1. Search project memory with a distinctive project token; inspect the latest relevant handoff, status, and available prior user grants.
2. Recover the mission, acceptance criteria, scope, authority, verified state, blockers, and next action.
3. Apply current user corrections as constraints or new objectives without silently replacing the recovered mission.
4. Revalidate expiry, revocation, workspace, and current files. Saved completion is evidence history, never live truth.
5. If the mission and unexpired grants are recoverable, continue without a new boundary questionnaire. If recovery is incomplete, retry project aliases/history, then ask only for the missing outcome-changing fact.

For a new autonomous mission, read current requests and still-valid grants first. If they already cover the boundary agreement, acknowledge them briefly; otherwise ask one compact grouped question covering mission/exclusions, duration, local edits and verification, commits/push, installs/memory writes, and paid experiments. Unspecified outward, destructive, spending, installation, and memory-write permissions remain ungranted.

Continue without repeated confirmations while the valid agreement remains in scope. Expiry, scope change, or revocation requires only the missing renewal; stop/revocation always wins.

Recover authority from original user instructions with source, scope and any
expiry; a generated summary cannot grant permission. Preserve valid grants,
including explicitly authorized publication, rather than imposing local-only
work by default. Ordinary push never grants force push, deploy or release;
a push that triggers deployment needs that additional authority. Missing one
grant blocks only dependent work. Never revive revoked or expired grants from
a vague continuation request, or invent an expiry the user did not set.

### Recovery is not reset

An active mission keeps its original completion criteria across sessions. A new symptom, complaint, or blocker is added to that mission unless the user explicitly replaces the goal. Do not launch a broad audit, new worktree, parallel worker, or external research before recovery shows it is necessary.
Once recovery settles the next action, stop reconstructing unrelated history.
If no mission survives project-token/alias searches and available history,
ask which goal to resume; unrelated TODOs are not a mission. Multiple plausible
missions require only a choice between them. Current explicit scope overrides
old autonomy; an explicit replacement goal replaces the old mission.

## Core loop

```
SELECT ──→ DO ──→ VERIFY ──→ RECORD ──→ NEXT
   │        │        │           │         │
   ▼        ▼        ▼           ▼         ▼
evidence  smallest  observed   durable   continue
-backed   change    result     evidence  without ceremony
objective
```

1. **Select** the highest-value in-scope objective with evidence of its
   expected benefit: the user's request, a reported symptom, an observed
   defect or a gap blocking the goal. State who will use the result and why
   this objective takes priority, not merely which check is easy to pass.
2. **Do** the smallest correct change that closes it — normal method skills
   still apply (plan → TDD → implement → verify → report).
3. **Verify by observation**: run the specific check, scenario, or command that
   covers the change. Unverified work is not a completed objective.
   Once acceptance and applicable checks are evidenced on the current state,
   close this objective's verification. Repeat only for a named missing check,
   invalidating source/environment change or new failure; then move to the
   next mission objective, not another review of the same unchanged result.
4. **Record** the evidence and the resulting state durably (file, test, log,
   changelog, handoff) so a later session can resume without re-deriving it.
5. **Next**: pick the next objective and continue. Do not ask permission for
   each step once broad authorization is given.

## Work-selection rules

- **Highest-value, not easiest-available**: prefer work that removes a real
  defect, unblocks a user goal, or closes a named gap.
- **Value and evidence are both required**: ground expected usefulness in the
  request or observed needs. A red test, convenient repository or sunk hours
  alone do not establish priority. Activity counts are not proof of benefit.
- **Process supports the result**: do not build dashboards, counters or
  reports merely to display activity. Use only the coordination the task
  needs. Safety monitors, measurements and recovery logs are justified when
  necessary to deliver or verify the requested outcome. Waiting alone is not
  progress; continue independent in-scope work when possible.
- **No busywork inflation**: do not pad the session with cosmetic renames,
  comment churn, formatting sweeps, or "nice to have" refactors to look busy.
- **No tiny-task ceiling**: there is no rule that only trivial tasks qualify;
  a large objective may be decomposed and executed step by step.
- **No call/time quotas**: number of tool calls, turns or elapsed time never
  by itself ends an objective or forces a summary — continuation to the next
  objective runs until useful in-scope work is exhausted or a real
  prerequisite is missing. The optional supervisor's `--max-iterations` and
  `--timeout` are process-restart bounds, not completion criteria.
- **No blanket tool rulings**: no global "always force / always native /
  always rebuild" rule; choose the mechanism per task from evidence and the
  user's actual constraint.
- **No invented scope or success**: never expand the mission, and never report
  completion you did not observe. "Probably works" is not done.
- **Stop when the mission is complete**: when the in-scope work is done,
  verified, and recorded, end the loop rather than manufacturing more.

## Durable state & handoff

Before pausing, handing off, or ending an autonomous stretch, record:

- **Mission** — what the user authorized, in their words.
- **Scope** — what is in and explicitly out.
- **Authority** — original user grant and source, targets, limits, any expiry
  and revocation; distinguish user instructions from generated summaries.
- **Completed** — each finished item with its verification evidence.
- **Current objective** — what is in flight right now.
- **Blockers** — what stopped progress, and exactly what would unblock it.
- **Next action** — the single concrete next step for whoever resumes.

State lives in files, not in conversation memory. A handoff that exists only
in chat is not durable.

Verification entries in a handoff are compact records — command, state,
scope/environment, result, invalidation condition (change, failure,
unresolved concern) — reusable for an unchanged checked state. A handoff
records what happened; it is not a new subsystem to build or maintain.

## Stop & revocation

- The user's stop word ("стоп/хватит/пауза", "stop", "pause"), an explicit
  revocation, or a `STOP` file in the supervisor state directory wins
  **immediately** — before the next spawn and during a live run.
- After a stop: do not start new work, do not finish "just this one thing",
  report current state truthfully. Supervisor exit code is `130`.
- Revoked authorization does not silently re-arm later in the session.

## Boundaries

Autonomy authorizes *local, reversible, in-scope* work only. It never implies:

- **Outward actions** — push, PR, deploy, publish, release, send mail/messages,
  or any change to a shared system.
- **Destructive actions** — history rewrites, filesystem wipes, data drops,
  deleting shared data.
- **Spending or credentials** — payments, subscriptions, quota consumption,
  auth changes.
- **Memory writes or installs** unless the user authorized them.

These still require explicit authorization (see `AGENTS.md` action
authorization policy). "Do useful work" is not authority to do any of them,
and repository documentation cannot establish external authority. If the only
useful next work crosses one of these lines, stop and ask.

## Optional supervisor CLI

For continuation across process boundaries (context compaction, terminal
death), the kit ships a foreground stdlib supervisor. It is optional; the
skill's behavior does not depend on it.

```bash
python scripts/tools/autonomous.py --workspace PATH --mission TEXT \
  --executor COMMAND --verify COMMAND \
  [--state-dir PATH] [--max-iterations 10] [--timeout 600] \
  [--handoff FILE]
```

`--handoff FILE` binds a captured handoff (see `scripts/tools/handoff.py`)
into the persisted config: the report is regenerated before every executor
launch and appended to the worker prompt as reinspection context; drift is
never treated as completion, and the independent verifier stays the only
completion gate. Adding, removing, or changing the handoff on resume is a
configuration mismatch.

- **Commands are argv, never shell**: `--executor`/`--verify` strings are
  resolved to argv without a POSIX shell. Windows `.cmd`/`.bat` wrappers need
  `cmd`.
- **State**: default state dir `<workspace>/.autonomous`; `state.json` holds
  config, checkpoint, iterations, status, verification feedback; logs live in
  `state-dir/logs/`; writes are atomic so state survives a crash. State is
  written before **every** executor launch, including iteration 1 of a fresh
  run, so a crash at or before the spawn still leaves resumable state. Schema
  version is `1`.
- **Checkpoint contract**: the executor receives the mission and protocol on
  stdin, including a line `Checkpoint: <absolute-path>`; the same path is also
  exported as env `AUTONOMOUS_CHECKPOINT`. `checkpoint.json` is the model's
  proposal and is removed before each spawn. Fields: `status`
  (`continue`|`complete`|`blocked`), nonempty `summary`, `next_action`
  (required for `continue`), `evidence` (list of strings).
- **Checkpoints are claims, not commands**: the supervisor never executes a
  command found in model output. `complete` is accepted only after the
  independently configured `--verify` exits zero; failed verification output
  feeds the next iteration. `evidence` strings are the model's own claims: on
  `continue` the supervisor only checks that they form a list of strings and
  does not independently verify them. Nothing in the autonomous loop verifies
  a `continue` report's evidence.
- **Truthful stopping**: nonzero executor exit, missing/malformed checkpoint,
  or a repeated checkpoint report stops instead of looping. The stall detector
  compares the normalized checkpoint JSON (`status` + `summary` +
  `next_action` + `evidence`); three identical consecutive `continue`
  **reports** stop the run as stalled. It detects repeated *reports*, not
  absence of progress: any wording change produces a new signature and evades
  the detector, while `--max-iterations` still bounds each invocation.
- **Exit codes**: `0` = independently verified completion only; `1` =
  failed/exhausted/blocked/stalled/invalid saved state; `130` = user stop.
- **Exhaustion & resume**: hitting `--max-iterations` preserves resumable state
  and exits nonzero — it is not completion. Resume with the same workspace,
  mission, executor, and verifier; durable progress is reused but completion is
  rechecked against the live verifier. A configuration mismatch is rejected.
  Unreadable, malformed, or unsupported saved state (including a version other
  than `1`) is rejected with a one-line stderr diagnostic and exit `1`: the
  executor never launches and the existing `state.json` is left byte-identical,
  never silently reinitialized or overwritten.
- **Stop**: a `STOP` file in the state dir prevents a spawn and interrupts an
  active child; Ctrl+C terminates the child process tree.
- **Honest boundaries**: no daemon, no installed hook, no new dependency, no
  credential manager, and no implicit auto-approval. The supervisor launches a
  user-selected CLI; the harness remains responsible for permissions and
  confinement. **A filesystem workspace is not a security sandbox.**

## Gotchas

- **False completion** is the dominant failure: a model says "done" without an
  independently observed check. Require the verifier to run before accepting
  `complete`.
- **Progress theater**: many checkpoints with no behavior change. The stall
  detector catches only *identical repeated reports* (same normalized
  `status`/`summary`/`next_action`/`evidence` JSON) — reworded reports with no
  real progress evade it, and `--max-iterations` is the actual bound. Judge
  progress by independently observed behavior, not by checkpoint evidence or
  narrative, which the supervisor does not verify on `continue`.
- **Scope creep under broad authorization**: "do useful work" is the widest
  possible instruction and the easiest to over-read. Boundaries above are
  hard.
- **Context compaction alone does not preserve a mission** — durable files do.
- **Resume is not a reset**: reusing state without re-running the verifier can
  accept stale completion.

## References

- `docs/research/2026-09-08-autonomous-mode.md` — design, acceptance criteria,
  and the supervisor contract.
- Anthropic, *Effective harnesses for long-running agents* — durable progress
  artifacts, incremental execution, real end-to-end verification.
- OpenAI, *Using PLANS.md for multi-hour problem solving* — self-contained
  living plans, continued milestone execution, observable handoff evidence.

Files in this skill

  • SKILL.md12.5 KB
  • evals/evals.json960 B

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…