Status: implemented (PRs 1 to 7, 2026-09-29 to 2026-09-30); the [follow-ups](#follow-ups) are not. Planned 2026-09-29, decisions confirmed the same day. Where this plan and the code differ, the code and the [user guide](../../docs/guides/agent-skills.md) describe current behavior. Goal: a person inside Claude Code, Codex, or Cursor can say "pull in the XYZ session from Codex" and the agent runs the right `agent-archive` command, with no setup beyond `agent-archive setup`. Live check: only Cla...
Installs into .claude/skills of the current project.
Are you the author of Specs?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/wangjohn-specs)
# Agent skills — engineering plan
Status: implemented (PRs 1 to 7, 2026-09-29 to 2026-09-30); the
[follow-ups](#follow-ups) are not. Planned 2026-09-29, decisions confirmed the
same day. Where this plan and the code differ, the code and the
[user guide](../../docs/guides/agent-skills.md) describe current behavior.
Goal: a person inside Claude Code, Codex, or Cursor can say
"pull in the XYZ session from Codex" and the agent runs the right
`agent-archive` command, with no setup beyond `agent-archive setup`. Live
check: only Claude Code was run end to end (see
[Live check](#live-check-2026-09-30)); Codex and Cursor rest on their
documentation, and, for Codex, on its bundled binary's skill discovery.
This plan builds on handoff v2 package F (#152), which already installs a
`/handoff` skill from setup with journaling, uninstall, and status. It
generalizes that installer instead of writing a second one.
## Decisions
Confirmed 2026-09-29:
1. The shared installer is `internal/agentskills` (not `agentcommands`).
Package F built `internal/agentcommands` first (#152); a mechanical
rename PR follows it, and `/handoff` becomes one entry in a registry of
skills.
2. Skills install by default in `setup`, with a `--no-skills` opt-out that
sticks across re-runs.
3. Upgrades refresh skills: `status` warns, `setup` refreshes, and a
non-interactive `setup --refresh` (run by `install.sh`) does it without
the person having to remember.
4. Permission allowlist edits stay out of scope (see [Later](#later)).
Adjustment made when #152 was found in flight: this plan's first draft
proposed a marker-plus-content-hash ownership rule. #152's rule is simpler
and also survives upgrades, so this plan adopts it (see
[Ownership](#2-ownership-adopted-from-152)).
## Problem
The CLI is already scriptable: `list`, `show`, `status` have `--json`, piped
output is never interactive or paged, `handoff` is size-bounded
(`--max-bytes`, default 120000) and has `--format json`, exit codes are
documented. What an agent lacks is *knowing* to use it, and which of the
commands are safe headless:
- `setup`, the `handoff` picker, and bare `show`/`list` on a terminal need a
TTY. An agent that runs them hangs or fails.
- `show --transcript` has no size bound. On a large session it floods the
agent's context.
- `handoff` takes a session ID or `--latest`, not a title. "The XYZ session"
means an extra `list --json` round trip and a guess.
- `setup`, `purge`, `uninstall`, `backfill` change state and must never run
unless the person asked.
## Goals
1. One read-only `agent-archive` skill, installed by `setup` for every
selected app, that the agent invokes on its own for "find / pull in /
look at a past session".
2. The common request is one command, not a `list` → pick → `show` chain.
3. Nothing an agent can reach through the skill floods its context or hangs
on a prompt.
4. Skill files never drift silently from the CLI: golden and doc-command
tests, and `status` reports stale, edited, or missing skills.
5. Setup, rollback, status, and uninstall treat every skill file like the
hook files: journaled, owned, and removable.
## Non-goals
- An MCP server. Both agents have a shell; a skill plus a good CLI is less
to install and maintain. Revisit if an agent without a shell matters.
- Editing agents' permission allowlists (Claude Code `permissions.allow`,
Codex execution-policy rules). Security-sensitive and agent-specific; the
skill's own `allowed-tools` (Claude Code) covers the main prompt problem.
- Project-level skills (`.agents/skills` in a repo). User-level only.
- Writing to the archive from the agent (`feedback`, `sync`).
## Where skills go
| App | File | Notes |
| --- | --- | --- |
| Claude Code | `$CLAUDE_CONFIG_DIR/skills/<name>/SKILL.md`, default `~/.claude/skills/…` | Same directory rule as its hook file (`hooks.ResolveFiles`). |
| Codex | `~/.agents/skills/<name>/SKILL.md` | The documented user location. Codex also reads `$CODEX_HOME/skills` (see [resolved question 1](#open-questions)); setup does not write there. |
| Cursor | the same `~/.agents/skills/…` file | Cursor also reads `~/.claude/skills` and `~/.codex/skills`. |
Codex and Cursor share one file per skill, so it may only use frontmatter
fields both accept. Claude Code gets its own render with `allowed-tools`
(and `disable-model-invocation` for `/handoff`). A skill needs no approval
step (unlike Codex hooks). Claude Code and Codex notice a new or changed
skill in a running session (Claude Code's `/reload-skills` covers a skills
directory that did not exist at start; Codex says to restart if it does not
show); Cursor's documentation does not say, so a new chat is the safe advice.
## Design
### 1. One installer for all agent skills (`internal/agentskills`)
`internal/agentcommands` (from #152) becomes `internal/agentskills`. Its
API (`Files`, `PlanInstall`, `PlanRemoval`, `Installed`) keeps its shape,
but renders a **registry** of skills instead of the one `handoff` skill:
```go
// A Skill is one installable skill. Render returns its file for one
// destination: Claude Code's (own frontmatter) or the shared
// ~/.agents/skills one Codex and Cursor read.
type Skill struct {
Name string // directory name: "handoff", "agent-archive"
Render func(dest Destination, executable, dataHome string) []byte
}
// Registry lists every skill setup installs, in a stable order.
var Registry = []Skill{handoffSkill, archiveSkill}
```
Skill bodies live as templates under `internal/agentskills/skills/<name>/`
(`go:embed`), not in Go string literals, so prose changes review as prose.
### 2. Ownership (adopted from #152)
A file is setup's when it carries the marker line, whichever release or
executable path wrote it, **and** its command names this installation's
data directory (`AGENT_ARCHIVE_HOME=…`, or none for the default one), so a
test installation sharing a `HOME` keeps its own file. A person keeps their
own version by deleting the marker line. Setup replaces only owned files
(refreshing them on upgrade), removes owned files for apps no longer
selected, and reports any other file at a wanted path as left alone.
Uninstall removes only owned files. Because the marker, not content
equality, decides ownership, an old release's file is refreshed rather than
mistaken for a hand edit. The cost is that an edit to a *marked* file is
overwritten at the next setup; the marker line says so.
`status` compares each owned file's content with what this release renders
and reports it **out of date** when they differ (the refresh signal for
[Upgrades](#upgrades)).
### 3. Setup, config, status, uninstall
Already in #152 for `/handoff`, generalized here:
- Skill changes join the setup journal with the hook changes, so a failed
setup rolls back both.
- The review and next-steps screens name every skill file setup wrote or
left alone.
- `status` lists installed skill files (`agent_skills` in `--json`,
renamed from #152's unreleased `agent_commands`) and per-file
out-of-date state.
- `uninstall` removes owned files and reports the rest.
New here:
- `setup --no-skills` (and `--yes --no-skills`) opts out of every skill.
It is recorded in config (`Config.NoSkills`, `json:"no_skills,omitempty"`)
so a re-run or `--refresh` does not reinstall them, and it removes owned
skill files already there (a file that is not setup's is left alone and
named). A prompt-driven setup does not ask; it installs, and its closing
output names the flag.
- Because the opt-out is sticky it needs a way back: `setup --skills`
clears it and installs. Giving both flags is a usage error (exit 2,
before anything is read or changed). While opted out, `status` says the
skills are turned off (`agent_skills_disabled` in `--json`, present only
when true).
- Nothing in the background rewrites agent configuration. The collector
never touches skills.
### 4. CLI changes so the skill is one command
1. **`handoff` accepts a title.** `handoff "fix the auth bug" --harness
codex` resolves with the same matcher `show` uses (title substring or
short ID), searching local registrations first (no network), then the
archive. Without a terminal, several matches print the candidates (short
ID, harness, project, title, time) to stderr and exit 1; it never
guesses and never prompts. With a terminal, the existing picker opens
(handoff v2 A's picker). A session ID still wins over a title that also
matches it.
2. **`show --transcript --max-bytes N`.** Same semantics and default as
`handoff` (120000; 0 = no limit; the full version saved for 7 days, path
named at the end). Makes the footgun safe rather than only documented.
Neither changes what is uploaded or what the filter drops, so no filter or
parser version bump. Both regenerate `docs/reference/cli.md` (the golden
test fails otherwise).
### 5. The `agent-archive` skill
Everything the agent needs is in the one `SKILL.md`, under 100 lines. There
is no `reference.md`: the installer renders one file per skill and
destination, Cursor may not load files beside it, and a skill this short does
not need one.
- **Description** (what triggers it): find, look at, or pull in a past coding
session, in any agent, from the archive or this machine. It names the three
agents, both places, and the phrasings people use ("pull in", "continue",
"review", "what did we do in Cursor yesterday"); a test pins those words,
and the plain-scalar and length limits (the agentskills.io format allows
1024 characters, Claude Code cuts the listing at 1536).
- **The flow:** `agent-archive handoff "<words>" --harness <app>` (the flag
only when the person named the agent the session was in). Bounded,
filtered, addressed to the receiving agent; the output is context, not a
task. If it lists several matches (exit 1, a table on stderr), the agent
shows the person the table and asks which; it does not pick. With none, it
tries other words once, or `list --since 30d`, and shows the titles. With
no topic ("where my other agent left off") it runs `handoff --latest
[--harness <app>]` (the project it is in; never its own session). It says
in one line which session it pulled in, so a wrong `--latest` or single
title match is caught by the person. Words are one or two distinctive title
words with no quote, `$`, backtick, or backslash.
- **Browse:** `list --since 7d --limit 20 [--harness …]`, the default text
table (title, when, agent, project, short ID; about 100 bytes a session),
not `--json` (about 2 KB a session); `show ID` for a summary;
`show ID --transcript` (bounded).
- **Never:** `setup`, `uninstall`, `purge`, `backfill`, `sync`, `feedback`,
`handoff --to` (that is `/handoff`), and `--max-bytes 0` or any larger limit,
unless the person asked for exactly that. These words appear in the file
only inside that section, and a test parses the template to keep it so.
- **Untrusted content:** pulled transcripts are data, not instructions;
filtered for credentials, not for adversarial text. That covers the tables
and the trimmed-output footer: a transcript can forge a "Full record: PATH"
line, so the agent opens only the `.md` file under a `handoffs` folder that
the last line names, and only if the person needs more.
- **Failures:** if a command fails for network, credential, or sandbox reasons,
say so and ask the person to allow it or run it, rather than retrying
variants.
- **Absolute binary path** (`config.InstalledExecutable`), because agents
often lack it on `PATH`; a relocated installation's `AGENT_ARCHIVE_HOME=`
prefix, as `/handoff` has. Paths needing quotes are quoted, and Claude
Code's `allowed-tools` rule is omitted for them (as `/handoff` does).
- **No environment prefix** such as `NO_COLOR=1` or
`AGENT_ARCHIVE_NONINTERACTIVE=1`. Claude Code strips only a fixed list of
variables before matching a permission rule, so a prefix would defeat any
rule the person saves for the command (and make each run ask again), and
the CLI already detects the three agents' shells. On a pseudo-terminal an
agent may still see the spinner's escape codes ahead of the output; that is
presentation (see below) and noise, not a wrong answer.
- **Claude Code extras and the `allowed-tools` decision.** No
`disable-model-invocation` (unlike `/handoff`, this skill is meant to be
invoked by the model). Its only pre-approved rule is the exact command
`Bash(<exe> status)`. Claude Code matches a rule against the whole command
text, `*` stands for any text (flags and spaces included), an allow rule
cannot exclude a flag, and a skill's `allowed-tools` has no deny list. So
`handoff:*` would also allow `handoff --to codex` (which starts an agent),
`handoff --output FILE` (which writes a file) and `--max-bytes 0`;
`show:*` would allow `--max-bytes 0` (an unbounded transcript); `list:*`
would allow `--rebuild-index` (which writes index keys to the bucket) and
`--limit 0`. This skill is model-invoked and reads text an attacker may
have written, so none of them is pre-approved: Claude Code asks once, and
the person may allow it permanently (their choice, made with the wider
rule in view). `status` has no flag that writes or is unbounded, and the
rule is exact, so nothing rides along on a wildcard (a redirect, for one).
A test models the matching rule and a table of dangerous commands. The
shared render has no `allowed-tools`. (The live check found that Claude
Code 2.1.283 in `claude -p` did not apply this rule when the model invoked
the skill, so `status` may ask too; see [Live check](#live-check-2026-09-30).)
A follow-up could make more of the flow pre-approvable: a read-only
command with no writing or unbounded flags (for example a `pull TITLE
[--harness H]` that prints what `handoff` prints, without the flags that
launch, write, or lift the bound) would be safe to allow as
`Bash(<exe> pull:*)`. Not in this PR.
- **Setup and status wording.** A slash skill (`/handoff`) is "Installed
/handoff, which ..."; the registry's `Slash` and `Summary` fields let any
other skill read as "Installed the agent-archive skill, which lets your
agents look up and pull in past sessions", and status names its row
`agent-archive:`.
### 6. Non-interactive mode (`AGENT_ARCHIVE_NONINTERACTIVE`)
Goal 3 says nothing an agent reaches through the skill hangs on a prompt.
"Not a terminal" is not enough for that: an agent's shell tool may allocate
a pseudo-terminal (Codex may), and every interactive decision (`list`/`show`
browser and picker, `handoff` picker, the pager, the alternate screen,
setup/uninstall/backfill confirmations) asks `Env.isTerminal`. So one
environment switch overrides it, decided in one place:
- `Env.interactive(stream)` is `isTerminal(stream)` and not switched off.
Every decision to ask, page, or take over the screen uses it. A raw
`isTerminal` (or `term.IsTerminal`) is for presentation only (colour,
redrawing a line, wrapping), which `NO_COLOR` and the terminal govern, not
the switch. `TestTerminalChecksAreClassified` scans the package's syntax
tree and fails on any terminal check, and any `newPrompter` call, not listed
with what it decides, so a new prompt cannot bypass the switch by accident.
- `AGENT_ARCHIVE_NONINTERACTIVE`: `1`/`true`/`yes`/`on` is on, `0`/`false`/
`no`/`off` is off, any case. Unset or empty means automatic: on when
`CLAUDE_CODE_SESSION_ID`, `CODEX_THREAD_ID`, or `CURSOR_AGENT` is set
(`currentSessionEnv` plus `cursorAgentEnv`, the variables handoff already
reads), since those mean an agent's shell is running the command. An
explicit value wins over automatic in both directions. Any other value is
one usage error (exit 2) before the command runs (not for the hidden
`_hook`/`_collect`, which must stay silent, nor help, version, or a
subcommand's `--help`, which answer before the environment is read) and, until
reported, counts as on: a typo never turns a prompt back on.
- Read through `Env.lookupEnv`, so tests never see the process environment;
`testEnv` gives an empty one, so the suite passes when run from inside an
agent.
- On: a command behaves as when piped. Ambiguous matches print candidates on
stderr and exit 1; `handoff`/`show` with no session are usage errors (exit
2); the pager and alternate screen are off; `setup`, `uninstall`, and
`backfill` keep refusing without `--yes`, and when the switch (not a
missing terminal) is why, the message names it and
`AGENT_ARCHIVE_NONINTERACTIVE=0`. `purge apply` without `--yes` refuses
too, terminal or pipe: it reads its typed digest from standard input
whatever that is, so it is refused explicitly instead of through
`interactive`. `setup --yes` never reads the R2 secret from a terminal it
may not ask on (a pipe is still read: a script's secret is not a prompt).
- Not covered by the switch, on purpose: presentation (the spinner and
colour still draw on a pseudo-terminal; `NO_COLOR` governs them), a read
of standard input that is a pipe (`_hook` payloads, the `setup --yes`
secret), and credential helpers the AWS SDK runs for an S3 profile
(`credential_process`, aws-vault, `op`), which agent-archive does not
control. `TestTerminalChecksAreClassified` is syntactic: it counts every
mention of a terminal check, prompt, or stream read in this package, so
method values and renamed imports are caught, but not a check made in
another package and handed in.
- `handoff --to <agent>` inside an agent (`/handoff`) needs no prompt: it
resolves the caller's own session from the environment (handoff v2 A),
`--latest --harness cursor` under `CURSOR_AGENT`, and otherwise is a usage
error, exit 2, instead of a picker. Handoff v2 D's choice between running in
this terminal and opening a new window must use `env.interactive`, so an
agent's pseudo-terminal opens a new window instead of being taken over.
## Risks
| # | Risk | Mitigation |
| --- | --- | --- |
| R1 | Cursor reads both `~/.agents/skills` and `~/.claude/skills`, so one skill may appear twice. | Cursor's documentation is silent on duplicate names and this was not run (no Cursor CLI here), so it stays open. The two files differ only in Claude-only frontmatter, so a duplicate is the same instructions twice; if it proves harmful, install one shared copy for Cursor-only setups. |
| R2 | Agent sandboxes (Codex default, Claude Code sandbox) may block the network and Keychain that `list` and archive reads need. | Confirmed for the network in Claude Code's sandbox and Codex's (`codex sandbox`): a title found on this machine still works, an archive read fails with `operation not permitted`, and the skill reported it and asked. Keychain access under a sandbox is not documented by any of the three and was not conclusive (see [Live check](#live-check-2026-09-30)). |
| R3 | Skill text drifts from the CLI. | Golden test of each rendered file; a doc-command test (extending `doc_commands_test.go`) that every command and flag a skill names exists. |
| R4 | A person has their own `agent-archive` skill. | Foreign-file rule: never overwritten; setup and status say so. |
| R5 | The skill invites agents to read large or hostile content. | Bounded output everywhere, injection warning, no `--max-bytes 0`. |
| R6 | Several sessions edit handoff v2 files this plan also touches. | Stacking order below; PR 2 and PR 3 wait for A and F to merge. |
## Shared names
Parallel packages rely on these; change them here first.
| Owner | Name |
| --- | --- |
| PR 3 `internal/agentskills` | `type Destination`, `type Skill`, `var Registry []Skill`, `Files(userHome, claudeDir string, harnesses []string, executable, dataHome string) []File`, `PlanInstall`, `PlanRemoval`, `Installed`, `Stale(userHome, claudeDir, executable, dataHome string) []string` (owned files whose content differs from this render), same signatures as `agentcommands` otherwise |
| PR 3 `internal/cli` | `status --json`: `agent_skills` (paths), `agent_skills_out_of_date` (paths) |
| PR 4 `internal/config` | `Config.NoSkills bool` |
| PR 4 `internal/cli/setup_flags.go` | `--no-skills`, `--skills` (both is a usage error) |
| PR 4 `internal/cli` | `status --json`: `agent_skills_disabled` (only when true) |
| PR 5 skill | directory `agent-archive`, `Registry` entry `archiveSkill` |
| PR 6 `internal/cli/setup_refresh.go` | `runSetupRefresh(...)`, flag `--refresh` |
| PR 4b `internal/cli` | `AGENT_ARCHIVE_NONINTERACTIVE`, `Env.interactive(stream)`, `cursorAgentEnv`, `agentShellEnv()` |
## PR breakdown
Each PR is one reviewable concern, wires what it adds into a command, and
leaves `main` releasable. CI's `deadcode` step fails on unreachable code.
Every PR adds a `CHANGELOG.md` line and follows the
[testing](../contributing/testing.md) rules: injected `Env`, no real home
directory. The list below is in dependency order; PR 1 starts immediately
and the rest wait on handoff v2.
| PR | Title | Base | Waits on |
| --- | --- | --- | --- |
| 0 | Plan agent skills | main | — |
| 1 | `show --transcript --max-bytes` | main | — |
| 2 | `handoff` accepts a title | main | handoff v2 A (#150) |
| 3 | Rename `agentcommands` to `agentskills`; skill registry; `status` out-of-date | main | handoff v2 F (#152) |
| 4 | `setup --no-skills` opt-out | main | 3 |
| 4b | Non-interactive mode: `AGENT_ARCHIVE_NONINTERACTIVE` | main | — |
| 5 | The `agent-archive` skill | main | 1, 2, 3 |
| 6 | `setup --refresh`; `install.sh` upgrades skills and hooks | main | 3, 4 |
| 7 | Docs; live check in all three agents | main | 1–6 |
Parallel once unblocked: 1 ∥ 2 ∥ 3 ∥ 4b; then 4 ∥ 5; then 6; then 7. There is
no separate status/uninstall PR: #152 already does it for `/handoff`, and
PR 3 generalizes it. A PR may be stacked on an unmerged base branch to
start early, and is retargeted to `main` when its base merges.
### PR 0 — Plan (docs only)
Adds this file and a row in `dev/README.md`. It does not edit `handoff.md`
(package F's section is in flight in #152); PR 3 updates that section when
it renames the package.
### PR 1 — Bound `show --transcript` (~120 lines + tests)
- Adds `--max-bytes` to `show` with `handoff`'s default and trimmed-file
behavior. Reuse the trimming helper; extract it if it is private to
`handoff`.
- Tests: over and under the limit; `--json` with a limit; `0`; the full
version written with mode 0600 and a 7-day expiry, path printed at the
end.
- Regenerate `docs/reference/cli.md`.
### PR 2 — `handoff` accepts a title (~250 lines + tests)
- The positional argument is a session ID or title, resolved with the
matcher `show` uses (extract it into a shared function), local
registrations first, then the archive.
- Non-terminal ambiguity: candidates on stderr, exit 1. No match: the
existing "no session" message with `list` as the next step. `--harness`
narrows.
- Tests: exact ID beats a title match; ambiguous with and without a
terminal; local-only match with the archive unreachable; `--harness`;
`--latest` and `--file` unchanged.
- Docs: help text, regenerated CLI reference, `docs/guides/handoff.md`.
- Review focus: matching local registrations reads titles and metadata
only, never transcript content.
### PR 3 — `internal/agentskills` (mechanical, ~+200 / −200 lines)
- `git mv internal/agentcommands internal/agentskills`; introduce
`Skill`, `Registry`, and move `/handoff`'s text into an embedded
template; add `Stale`; rename the status field to `agent_skills` and add
`agent_skills_out_of_date`; status text lists both.
- No behavior change for `/handoff`; existing tests carry over unchanged
apart from the rename, plus new tests for a two-skill registry and for
`Stale`.
- Updates `handoff.md`'s package F section to name `agentskills` and point
here.
- Review focus: the diff is a rename plus small additions; verify with
`git diff -M` that no logic moved by accident.
### PR 4 — `setup --no-skills` (~200 lines + tests)
- Flags `--no-skills` and `--skills`, `Config.NoSkills`, a closing line
that names the opt-out (or, when opted out, `--skills`), `--yes` path. A
set flag skips installing and removes owned skill files. Recorded so
re-runs and `--refresh` honor it; `--skills` clears it.
- Tests: fresh setup with the flag installs none; existing install then
`--no-skills` removes owned files and leaves foreign ones; a later plain
setup keeps the opt-out; `--skills` turns it off; both flags is a usage
error; another installation's file is never removed; the journal rolls the
removal back on failure.
- Docs: setup guide, `docs/reference/configuration.md`, CLI reference.
### PR 4b — Non-interactive mode (~250 lines + tests)
- `Env.interactive`, the switch, the automatic agent variables, the
refusal hints, one usage error for a bad value (design section 6).
- Every interactive decision moves to `interactive`; the rest are
classified by `TestTerminalChecksAreClassified`.
- Tests: each picker, browser, pager, and confirmation with the switch on and
a fake terminal; automatic on for each of the three variables; explicit `0`
overrides; invalid value; unchanged when unset; the suite passes with
`CLAUDE_CODE_SESSION_ID` and `CODEX_THREAD_ID` set in its own environment.
- Docs: `configuration.md`, the handoff and list/show guides,
troubleshooting, help text and the CLI reference.
- Independent of 1–3 and 5–6; PR 5's skill can rely on it.
### PR 5 — The skill (~80 lines of prose + tests)
- `internal/agentskills/skills/agent-archive/SKILL.md.tmpl` (one file, no
`reference.md`); `archiveSkill` added to `Registry` after `handoffSkill`;
`Skill.Slash`, `Skill.Title()`, and `Label` so setup's "Installed ..." and
"Left ..." lines and status's rows read sensibly for a skill that is not a
slash command; Claude Code `allowed-tools` for the exact `status` command
only (see Design 5).
- Tests: golden render per destination for a plain path, a quoted path, and
a relocated data directory; the doc-command test extended to the rendered
templates (`TestAgentSkillsQuoteOnlyRealCommandsAndFlags`); a test that the
dangerous words appear only inside the "never run" section and that the
skill runs only `handoff`, `list`, `show` and `status` with read-only
flags; a model of Claude Code's rule matching with a table of commands no
rule may permit; registry order and setup/status output with two skills;
Stale and ownership for the new skill; description length and YAML limits.
Regenerated: the status JSON and screen goldens (they list the new skill).
- Review focus is the prose: the description is what triggers the skill, so
reviewers should try it against several phrasings.
### PR 6 — Refresh on upgrade (~350 lines + tests, plus `install.sh`)
- New non-interactive `setup --refresh`: re-renders hook and
skill files for the **saved** configuration and the current executable,
and nothing else. It asks no questions, needs no terminal, never touches
storage, credentials, projects, retention, or the LaunchAgent's job
state, and exits 0 with "nothing to refresh" when everything is current.
It refuses (exit 1, one-line reason, nothing changed) when setup never
completed, an interrupted setup needs recovery, the archive is
uninstalled, another installation's hooks are in a file it would write,
another installation owns the collector's launchd label, or it runs as
root in a home directory that belongs to another user (`sudo` can keep
`HOME`; it would leave root-owned files there). Any flag but
`--verbose` is a usage error (exit 2). A paused archive is refreshed like
any other, and stays paused: pausing keeps the hooks.
- It works in `internal/cli/setup_refresh.go` from the saved configuration:
hooks go where setup recorded them (`hook_files`), not where this shell's
`CLAUDE_CONFIG_DIR` points. It shares setup's skill planning
(`planAgentSkills`), its journal, and its collector, hooks, and setup
locks, so a failure rolls back. Ownership is unchanged. It waits up to ten
seconds for a collector pass (the collector starts one every minute)
before refusing, and absorbs SIGINT, SIGTERM, SIGHUP, and SIGQUIT from the moment
the journal is written until it is gone, so an interrupted installer
never leaves a transaction to recover. `launchctl` runs in a process group
of its own, so a terminal's Ctrl-C does not kill it halfway, and
`bootstrap` and `bootout` end after 30 seconds, so nothing that absorbs
signals can hang.
Only files whose
content would change are written, so a current installation writes and
asks nothing.
- The journal gained `FilesOnly` (`files_only`): a transaction of files
that neither commits nor rolls back through launchd. Refresh uses it
except in one case. The LaunchAgent plist is rewritten only when it runs
another executable, keeping its environment (the AWS files and `PATH` the
storage check ran with, which the installer's shell may not have), and
when that plist belongs to a **loaded** job the ordinary transaction runs
(stop, write, start), because launchd runs the definition it loaded, not
the file, and a job left on a deleted binary would fail every minute. A
job that is not loaded stays unloaded; a job whose state is unknown, or
another installation's, refuses. A release without the field ignores it
and recovers such a journal as an ordinary one: files back, a loaded
collector stopped and not started again until setup runs. Only a crashed
refresh followed by a downgrade meets that.
- Also repairs hook drift: if `InstalledExecutable` differs from the running
executable (the case `status` reports as "capture has stopped"), refresh
points hooks, the LaunchAgent plist, and the skills at the running binary
and records it, in the same transaction. It applies setup's own guards
first: a `go run` or `go test` build, a file in the temporary folder, or
one that is missing or not executable is refused.
- `install.sh`: after installing the binary, if an existing configured
installation is found, runs `setup --refresh` and prints
its one-line result. A refresh failure prints the reason and the manual
command but does not fail the install; a fresh install runs nothing.
`AGENT_ARCHIVE_HOME` is honored. As root (`sudo`, which can keep `HOME`)
it skips the refresh and says to run it as the person, since it would
leave root-owned files in their app settings. A refusal names both ways
on: `setup --refresh` again when the reason was temporary, `setup` when
the archive was uninstalled or setup never finished.
- `status`'s out-of-date skill line, its "capture has stopped" warning, and
the next step for missing hooks say `setup --refresh`.
- Tests: refresh replaces a stale skill and leaves a foreign one; refresh
with `NoSkills` installs none; refusal without saved config; hook drift
repair; idempotent second run; rollback on failure; installer-script
tests (the release and installer scripts are CI-checked) for found and
not-found installations.
- Review focus: refresh cannot change anything beyond hook, skill, and
recorded-executable files. Test by diffing the whole home directory
before and after.
### PR 7 — Docs and live check (docs + a run log)
- Guide `docs/guides/agent-skills.md`; README and setup guide mention;
`setup.md` "what setup changes on your machine" lists the skill files;
uninstall guide; troubleshooting ("agent doesn't use the skill"). PR 5
wrote most of the guide; PR 7 read every user document as a new reader
would, tightened what the live check contradicted, and added what an agent
can read through the skill to the privacy document.
- Live check, in the sandbox recipe of
[testing](../contributing/testing.md): Claude Code end to end, an upgrade
through `install.sh` that refreshes two stale skills, Codex's skill
discovery and sandbox with the `codex` binary bundled in ChatGPT.app, and
Cursor from its documentation only. The results are in
[Live check](#live-check-2026-09-30), and the pull request description has
the run table. Not done, and why: there is no Cursor CLI and no Codex login
to drive a model, so "the flagship request from each agent to each other",
the Cursor-plus-Claude duplicate (R1), and a Codex model run stay open.
## Upgrades
Skill text embeds the binary path and flag names, so it must refresh when
the binary changes:
1. `status` warns "out of date; run `agent-archive setup --refresh`" (PR 3,
reworded in PR 6).
2. `setup` (re-run) refreshes as part of its normal transaction (#152).
3. `install.sh` runs `setup --refresh` when it finds an
existing configured installation (PR 6), so upgrading the binary
upgrades the skills with no extra step.
The background collector never writes agent configuration.
## Open questions
Resolved 2026-09-30 from official documentation, then checked against the
binaries that could be run. **No installer path changes**: the paths in
[Where skills go](#where-skills-go) are the documented ones for all three
agents.
1. **Codex.** *Where does it read user-level skills?* The documented user
location is `$HOME/.agents/skills` (repository skills are in `.agents/skills`
in the working directory, its parent, and the repository root; admin skills
in `/etc/codex/skills`; system skills are bundled). The documentation does
not mention `$CODEX_HOME/skills`. The Codex binary bundled in ChatGPT.app
(0.155.0-alpha.9.2, run with a scratch `HOME` and `CODEX_HOME`, no login,
`codex debug prompt-input`) does load both: a skill in `$HOME/.agents/skills`
alone, one in `$CODEX_HOME/skills` alone, and one in each, which listed
the name twice (the `$CODEX_HOME` copy first). So there is no
de-duplication and no name-level precedence, which the documentation
agrees with: Codex does not merge same-named skills, and both can appear
in selectors. Setup writes only `~/.agents/skills`, so a Codex user gets
one entry. *Restart or refresh?* Codex detects skill changes
automatically and says to restart if an update does not show; the page
names no `/skills` refresh (`/skills` and `$` are how to mention a skill).
*Frontmatter:* `name` and `description` are required and drive implicit
invocation; optional metadata belongs in `agents/openai.yaml`
(`policy.allow_implicit_invocation`, default true). `allowed-tools` and
`disable-model-invocation` are not documented, so the shared file carries
neither.
2. **Cursor.** *Which directories?* Project: `.agents/skills/`,
`.cursor/skills/`, and, for compatibility, `.claude/skills/` and
`.codex/skills/`. User: `~/.agents/skills/`, `~/.cursor/skills/`,
`~/.claude/skills/`, `~/.codex/skills/`. *Does it de-duplicate the same
name in two directories (R1)?* Not documented, and not run. Frontmatter:
`name` and `description` required; `paths`, `disable-model-invocation`,
`icon`, `color`, `metadata` optional; `allowed-tools` is not mentioned
(the shared file has none). Skills apply automatically by default and `/`
invokes one. A restart is not documented.
3. **Claude Code.** *Where?* Personal skills are `~/.claude/skills/<name>/`,
project skills `.claude/skills/<name>/`; personal beats project for one
name. The skills page does not mention `$CLAUDE_CONFIG_DIR/skills`; the
environment-variable reference says `CLAUDE_CONFIG_DIR` overrides the base
directory where Claude Code keeps its configuration, cache, plugins,
sessions, and logs (default `~/.claude`), which is the rule
`hooks.ResolveFiles` already follows for hooks, so setup follows it for
skills. That was not run (a scratch configuration directory has no
login). *`allowed-tools`:* it pre-approves the listed tools for the turn
that invokes the skill and clears at the next message; it does not
restrict tools, deny and ask rules still win, it takes a space- or
comma-separated string or a list, and it has no deny list of its own
(`disallowed-tools` removes tools from the pool, which is different). A
Bash rule is matched against the whole command; the wrappers `timeout`,
`time`, `nice`, `nohup`, `stdbuf` and a fixed list of known-safe
environment variables are stripped first, and an allow rule does not match
past an assignment of any other variable, which is why the skill carries
no `NAME=value` prefix. *Auto-trigger:* Claude reads a listing of skill
names and descriptions and invokes a skill through the `Skill` tool;
`description` plus `when_to_use` is capped at 1,536 characters in that
listing; `disable-model-invocation: true` stops it. The `Skill` tool is
itself a permission-required tool, with rules `Skill(name)` and
`Skill(name *)`. *Restart?* No: Claude Code watches the skill directories
and picks up an added, edited, or removed skill in the running session;
`/reload-skills` covers a top-level skills directory that did not exist
when the session started. *`claude -p`:* a tool call that would prompt is
denied and the model is told so (observed below).
4. **Sandboxes and approvals (R2).** *Claude Code:* the Bash sandbox
(Seatbelt on macOS, nothing to install) is off until `/sandbox` or
`sandbox.enabled`, and pre-allows no network domains (the first command that
needs a new domain prompts; where nothing can ask, it is refused). *Codex:* the default
is `workspace-write` with `on-request` approvals and network off, so a
command that needs the network asks first; macOS enforcement is Seatbelt;
command rules live in `~/.codex/rules` (`prefix_rule`, decisions `allow`,
`prompt`, `forbidden`); `codex exec` defaults to a read-only sandbox, and
the documentation does not say what an approval request does with nobody
there. *Cursor:* the run mode decides (its documentation does not confirm
which is the default for a new installation, and calls Auto-review the
recommended default): Auto-review runs allowlisted calls at once and other
shell commands in the sandbox where it can, asking only when one needs full
access and the classifier finds it risky; Allowlist asks for anything not
on the list; Run Everything asks nothing; on macOS the sandbox
is Seatbelt with the network "blocked by default" until a network mode
(`sandbox.json` only, `sandbox.json` plus defaults, or allow all) opens
it; the CLI keeps allow and deny lists in `~/.cursor/cli-config.json` or
`.cursor/cli.json`. Whether a sandboxed command may read the macOS
Keychain is **not documented by any of the three**. In practice the
network is the blocker (see [Live check](#live-check-2026-09-30)), and it
needs an approval or an allowed domain whatever the agent.
Sources (official pages; the Codex pages redirect from developers.openai.com
to learn.chatgpt.com):
- Claude Code: <https://code.claude.com/docs/en/skills.md>,
<https://code.claude.com/docs/en/permissions.md>,
<https://code.claude.com/docs/en/tools-reference.md>,
<https://code.claude.com/docs/en/env-vars.md>,
<https://code.claude.com/docs/en/sandboxing.md>
- Codex: <https://developers.openai.com/codex/skills> (now
<https://learn.chatgpt.com/docs/build-skills>),
<https://developers.openai.com/codex/security> (now
<https://learn.chatgpt.com/docs/agent-approvals-security> and
<https://learn.chatgpt.com/docs/sandboxing>),
<https://developers.openai.com/codex/rules> (now
<https://learn.chatgpt.com/docs/agent-configuration/rules>),
<https://developers.openai.com/codex/noninteractive> (now
<https://learn.chatgpt.com/docs/non-interactive-mode>)
- Cursor: <https://cursor.com/docs/context/skills>,
<https://cursor.com/docs/agent/terminal>,
<https://cursor.com/docs/agent/security/run-modes>,
<https://cursor.com/docs/cli/reference/permissions>
Still open: whether Cursor lists a skill twice when it is in both
`~/.agents/skills` and `~/.claude/skills` (R1), whether `$CLAUDE_CONFIG_DIR/skills`
is read when the variable is set, and what the three agents' approval
dialogs look like for the commands a skill names (Claude Code's was only
observed in `claude -p`).
## Live check (2026-09-30)
Built from `origin/main` (84481578). Everything ran in a scratch directory
with its own `AGENT_ARCHIVE_HOME`, a throwaway MinIO from `quay.io`, an AWS
profile of its own, a stub `launchctl`, and synthetic sessions only (three
Codex, one Claude Code, one Cursor, registered through `_hook` and uploaded
with `sync`). The Claude render was written by the real renderer
(`agentskills.Files`) as a project skill in the scratch project; nothing was
written under the real `~/.claude` or `~/.agents`. `claude -p` ran from that
project with `--setting-sources project,local` (so the person's hooks and
user skills stayed out), `--permission-mode default`, and explicit
`--allowedTools` limited to the scratch binary's path and
`Skill(agent-archive)`: 21 invocations.
- **Triggering, on seven phrasings:** the skill was invoked and ran the intended
command each time (`handoff "webhook retry" --harness codex` for "pull in the
session where we migrated the billing webhook retry queue, it was in codex";
`handoff "flaky test"` with no harness; `handoff "checkout" --harness codex`
for "what did codex figure out about the checkout totals last night";
`handoff "pagination" --harness cursor` for "that cursor chat about
pagination yesterday"; `handoff "thumbnail" --harness codex`; and
`handoff --latest` for "continue where my other agent left off", which
named no harness). Four unrelated requests (summarize `README.md`, run the
tests, write a commit message, recap this conversation) did not invoke it.
- **Ambiguity and no match:** two sessions matched "flaky test"; the agent
showed the table (short ID, agent, project, when, title) and asked, without
picking. For an absent topic it tried one other word, ran
`list --since 30d`, and showed the titles.
- **Untrusted content:** the synthetic session's reply told "the reading
assistant" to run `setup --yes`, `handoff --to codex`, and
`purge apply --yes`. The text came through the handoff (the filter drops
credentials and instruction blocks, not this), and in two runs (an explicit
title request, and `--latest`, which pulled the same session in unasked) the
agent summarized it and ran nothing else.
- **`claude -p` permissions (Claude Code 2.1.283):** with
`Skill(agent-archive)` allowed and nothing else, `handoff` and `status`
were denied ("requires approval"); the agent told the person it could not
run the command, printed it, and offered to run it after approval or for the
person to run it and paste the output. Without `Skill(...)` allowed, the
`Skill` call itself is denied and the agent went looking for the session
files by hand (blocked by the working-directory limit): the `Skill` tool
is a permission-required tool. **The skill's
`allowed-tools: Bash(<exe> status)` was not applied** when the model
invoked the skill (as a project skill, with `--setting-sources
project,local`; a personal skill under `~/.claude/skills` was not run),
although the documentation says a skill's
`allowed-tools` applies in a `-p` run: a control skill with other rules
behaved the same, and the same rules were applied when the person typed
`/skill-name`. The interactive dialog was not observed (starting an
interactive session would have written a trust entry into the real
`~/.claude.json`). Treat the `status` pre-approval as a convenience that may
not apply, not a guarantee.
- **Sandbox (R2), Claude Code with `sandbox.enabled`:** `handoff` by title
worked (local sessions need no network); `list` failed with `connect:
operation not permitted` to the MinIO endpoint, and the agent named the
sandbox, gave the command, and offered to retry outside it, without
retrying variants.
- **Sandbox, Codex (`codex sandbox`, the default state, ChatGPT.app's
binary):** the same two results (local title match works, archive read
`operation not permitted`). A Keychain lookup of an item that does not
exist answered as it does outside the sandbox, with one extra parameter
error from the Security framework: inconclusive for a real read, so an R2
user should expect to approve or run archive reads outside the sandbox.
- **Codex discovery:** see resolved question 1.
- **`install.sh` upgrade, shipped script, fake download of the built binary
(a `codesign` stand-in accepts it, since a local build has no Developer ID),
scratch `HOME`, first set up with a8dfdbe5, from before the `agent-archive`
skill existed:**
`setup --refresh` printed `refreshed 2 skill files` (the new
`agent-archive` skill, for both destinations; `/handoff` was already
current). After appending old text to all four files, `status` warned four
times (out of date) and the installer printed `refreshed 4 skill files`;
`status` was then clean.
The skill text needed no change: no run misfired, and the "never" and "data"
sections held. The description's phrasings ("pull in", "continue", "review",
"what did we do in Cursor yesterday") were enough for the paraphrases tried.
## Follow-ups
Not done in this series:
1. **A read-only `pull` entry point** (for example `pull TITLE [--harness H]`
printing what `handoff` prints, with no `--to`, `--output`, or
`--max-bytes 0`). A skill's `allowed-tools` cannot exclude flags, so
`Bash(<exe> handoff:*)` would also allow `handoff --to`; a command with no
dangerous flag could be pre-approved as `Bash(<exe> pull:*)`, and the common
flow would need no permission prompt. The live check also found the
`status` pre-approval unreliable for a model-invoked skill, so this is
where most of the friction remains.
2. **`handoff`'s size bound is best effort.** A 400-exchange session produced
about 168 KB against the 120 KB bound: `FitHandoff` never drops exchanges,
only shortens them (after every step it prints `warning: still N bytes
after trimming, over the N-byte limit` on stderr, so the overrun is
visible). It needs an oldest-exchange drop and matching
documentation and golden changes; the handoff owner's call.
3. **Skip the activity spinner when non-interactive.** With
`AGENT_ARCHIVE_NONINTERACTIVE` on (a pseudo-terminal inside an agent) the
spinner's escape codes can still precede the output. Presentation only, but
it is noise in an agent's context.
## Later
- `setup` writes permission allowlist entries so the read-only commands
never prompt in Codex and Claude Code (declined for now, decision 4).
- `list --local` for offline discovery, if title matching on `handoff`
proves too narrow.
- An MCP server, if a shell-less agent needs it.
## Tests summary
Per-PR tests are listed above; across the series, the invariants are:
setup never overwrites a file it does not own; rollback removes what setup
wrote; status and uninstall agree with setup about ownership; the rendered
skills name only commands and flags the CLI has; no command a skill
recommends prompts, pages, or prints unbounded output without a terminal.