Skip to content
Back to skills

Keep Going

ASecurity

Recover and continue after an interruption, rate limit, crash, or gap, or check on off-thread work that looks stalled: inspect its real output, act on evidence (resume, rerun, kill-and-restart), then continue the main task. Use when: asked to keep going, continue, or resume; 'what were you doing'; 'check the monitor', 'is it stuck', 'stop staring at it'. Gates killing or re-firing side-effectful work. To retire finished work and reconcile the ledger instead, use /session-flow:reconcile.

  • 13 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 2, 2026
ai-agentspythonrustgoshell

Security analysis

A100/100

Pro scans all 6 files and shows the line behind each finding

Scanned October 4, 2026

npx -y skills add melodic-software/claude-code-plugins --skill keep-going --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Keep Going?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Keep Going
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/melodic-software-keep-going/badge)](https://www.skillsdirectory.com/skills/melodic-software-keep-going)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
description: "Recover and continue after an interruption, rate limit, crash, or gap, or check on off-thread work that looks stalled: inspect its real output, act on evidence (resume, rerun, kill-and-restart), then continue the main task. Use when: asked to keep going, continue, or resume; 'what were you doing'; 'check the monitor', 'is it stuck', 'stop staring at it'. Gates killing or re-firing side-effectful work. To retire finished work and reconcile the ledger instead, use /session-flow:reconcile."
user-invocable: true
disable-model-invocation: false
metadata:
  workflow-stage: session
  summary: Recover after an interruption and continue where work stood
---

# Keep going

## Purpose

After an interruption, a rate limit, a crash, a disconnect, or just a
long gap, work started off the main thread may be paused, dead, or
silently finished, and the main task's position is easy to misremember.
The same recovery also applies live, mid-session, when off-thread work
*looks* stalled and someone asks you to check on it. This skill recovers
or verifies the off-thread work from its **real** state, reconciles the
main thread, and continues. Scope is general: the recovery is the same
whatever caused the interruption, so the cause is not diagnosed here.

Two failure modes this skill exists to prevent: **summarizing and
stalling** once a usage limit has lifted instead of continuing, and
**killing live-but-slow work on a hunch** instead of on evidence.

Where `/session-flow:handoff` deliberately pauses a session, keep-going is
the resume counterpart. It picks the work back up after any pause, planned
or not.

## Two ways in, same flow

1. **After an interruption**, a rate limit, crash, disconnect, or gap
   left work in an unknown state.
2. **A live-session poke**. Nothing was necessarily interrupted, but
   off-thread work looks stalled and you are asked to check ("check the
   monitor", "poke it", "is it stuck", "stop staring at it"). Verify
   against real output *before* acting; it may be alive and progressing.

Either way the flow is identical: inventory → inspect the real state →
align on the recorded goal → act on evidence → continue.

## Intent comes from the conversation, not the arguments

Infer *what* to resume or check from the conversation and the real
off-thread state, not from an argument. Arguments only narrow the target;
their absence never blocks. When the conversation makes the referent
clear, do not stop to ask "what should I keep going on?"; that stall is
itself the thing this skill removes.

## Steps

1. **Inventory off-thread work.** Enumerate everything running outside
   this thread, per the off-thread kinds in
   [`${CLAUDE_PLUGIN_ROOT}/reference/off-thread-work.md`](${CLAUDE_PLUGIN_ROOT}/reference/off-thread-work.md)
   an open-ended set (background tasks, shells, monitors, scheduled
   jobs, dynamic workflows, subagents), not a fixed catalogue.
2. **Inspect real state, never assume.** Read each item's actual state
   from the source of truth, per that doc's inspect-real-state invariant:
   do not infer "it probably finished" or "it probably died". Only the
   artifact tells you which.
3. **Goal alignment, before any recovery ACTION.** When the resume
   follows a `/session-flow:handoff`, read that file rather than trusting
   memory, and when the handoff's path was lost (a `/clear` without
   copying the resume prompt), recover it first with
   `/session-flow:find-handoff`. Read its `Original goal` section, then
   test the planned next actions against it: **say in one sentence how
   the next action serves that goal.** If you cannot, that is drift, not
   a wording problem. Stop, and either re-derive an action that does
   serve the goal or ask the user whether the goal changed. **A handoff
   carrying no `Original goal` is itself a defect:** do not infer the goal
   from the process the file describes, which is the thing that drifted,
   ask the user for it in their own words before continuing, and carry
   their answer into the next save-point.

   This check sits between inspection and action deliberately: steps 1-2
   only read, but step 4 resumes and restarts work, and restarting work
   that serves a drifted goal re-arms the drift before anything has
   tested it. One sentence, not a new stage. Its cost is nothing and its
   absence is the only signal that many sessions of faithful execution
   were aimed at the wrong thing.
4. **Recover per item. Act on evidence.** Classify against the real
   output and act:
   - **Progressing** (even if slow) → leave it; report it is alive and
     moving. Do not kill work that is making progress.
   - **Resumable** → resume it. Prefer a real resume over a restart when
     the mechanism supports one (a workflow resume reuses the cached
     prefix instead of redoing work).
   - **Dead but safe to redo** → restart it (subject to the autonomy
     policy below).
   - **Unrecoverable** → surface it plainly; do not fake a recovery.
5. **Reconcile the main thread.** Restate where the primary task
   actually stood, grounded in a fresh read of any plan / checklist /
   task artifact backing it, not a prior turn's claim, and continue it.
6. **Report.** Lead with anything waiting on the user (a gated kill or
   re-fire, a goal question), then one list: recovered, restarted,
   still-running, and lost / unrecoverable.

## Zone input (presence-gated, conservative)

Step 5 continues the main task in **this** session, and that is only right
when this session's context is still fit for the work. When the
`context-guard` plugin is enabled, resolve this session's zone word per
its reader contract before continuing (the contract owns the snapshot
path, staleness rule, and bands. Read them there; this skill consumes
only the resulting word and inlines no band values). Never substitute
your own estimate of the remaining window for the instrument's reading,
a resumed session's sense of its own budget is exactly the guess the
instrument exists to replace. Plugin not enabled, absent snapshot, or
`unknown` → judge from response quality alone, conservatively. A degraded
zone, or context-guard's evidence-degraded marker for this session, does
not stop the recovery (steps 1-4 are reads and evidence-gated actions
either way); it changes where the *continuation* goes. Route it with
`/session-flow:workflow continue`, whose router in
[`../workflow/context/continuation.md`](../workflow/context/continuation.md)
makes that choice.

## Active-verification protocol. Evidence before action

For any "is it stuck / check the monitor / poke it":

- **Read the real output first**. Monitor status, subagent
  transcript/output, task output, shell logs. Judge from the artifact,
  never from "it has been a while."
- **Progress-vs-elapsed is a suspicion-raiser only.** Slow relative to
  elapsed time tells you to look closer; it never by itself authorizes a
  kill. Confirm "dead" or "stuck" against the actual output.
- **When the evidence is ambiguous, treat the work as alive.** Killing
  live-but-slow work that was actually progressing is the failure mode to
  guard against.

## Autonomy policy. Resume freely, gate side effects

- **Auto-resume** safe, idempotent, read-only, or clearly incomplete
  work without asking. Recovery should not stall on confirmation for work
  that cannot double-fire. Keep going when a step needs no input, with
  status notes in the same message as the next action; never end on an
  offer to continue, a summary naming the next step without taking it, or
  a list of options that block nothing.
- **GATE** before RE-FIRING anything with external side effects, a push,
  a PR comment, a sent message, a deploy, a mutation, **and** before
  KILLING or RESTARTING off-thread work whose death you cannot prove from
  step 2. When the inspection cannot prove the action did NOT already
  land, or cannot prove the work is actually dead, stop and ask.
  Double-firing a side effect, or killing live work, is worse than
  pausing.

## After a usage limit lifts. GO, don't summarize

- If you are executing again, the block is already over: **continue the
  work**. Do not produce a summary and stop. Summarize-and-stall is the
  failure mode. Put status notes in the same message as the next action.
- The time-vs-reset check belongs to the **orchestration** case: when
  step 2 inspects a worker or subagent that is itself limited, compare the
  current time against the reset its limit message states, to decide
  whether that worker can proceed now. Run the bundled checker instead of
  doing this arithmetic yourself:

  ```shell
  python3 "${CLAUDE_PLUGIN_ROOT}/skills/keep-going/scripts/check-usage-limit-reset.py" "<limit message text>" --received "<ISO-8601 time the message appeared, with offset or Z>"
  ```

  Pass `--received` with the time the limit message appeared: the timestamp
  of the transcript entry that carries it, else the time you captured it. A
  message states only a clock time, so without `--received` a reset time
  already past on today's clock reads as passed even when the message meant
  tomorrow. When no time is known, omit the flag and treat exit `0` as
  provisional, like exit `1`.

  Exit `0` means the reset has passed, treat the worker as resumable now.
  Exit `1` means the limit is provisional until a live re-check of the
  current account confirms it (next bullet); only then hand back by invoking
  `/session-flow:handoff` via the Skill tool and stop. Exit `2` means the checker could not parse its input; read stderr
  first. `no reset clause found` means the message had none: say so plainly
  and ask the operator rather than guessing. `--received needs an offset or Z`
  means your own `--received` value was malformed: fix it and rerun. Exit `3` means the
  reset clause parsed but the IANA timezone could not be resolved (rare when
  the bundled `tzdata` under `scripts/vendor` is present); report the timezone
  failure rather than treating the message as unparsable. In a single
  interactive session, if you are running, the answer is already GO. A date-bearing
  form such as `resets Sep 8, 6pm (America/New_York)` is unparsed (exit `2`); never
  treat exit `2` as lifted.
- The limit **message text** (e.g. `resets 3:45pm`) is a capture bound to the
  account that emitted it. Two live readings of the *current* account count:
  the `/usage` view, which the operator relays because the model cannot open
  it, and the statusline `rate_limits` object, which this skill reads only when
  a statusline or hook exposes it to the session. For which fields that object
  carries and who receives it, see
  [statusline: Available data](https://code.claude.com/docs/en/statusline#available-data);
  the record below holds its as-of date and recheck trigger. This skill has no
  in-session account-identity signal, so a captured message never drives a
  still-blocked verdict by itself: re-check live before handing back. When no
  live reading is obtainable (headless, subagent, cloud, no statusline
  producer), ask the operator which account is active and whether it has
  headroom, as in the exit `2` path. Never invent a window and never conclude
  still-blocked.

  | Decision | Pointer | As of | Recheck trigger |
  |---|---|---|---|
  | A captured usage-limit message is account-bound. Still-blocked requires a live re-check of the current account, or the operator's answer when none is obtainable. The date-bearing `Sep 8, 6pm` form is unparsed (exit 2). | For the reset time a limit message carries, see [costs: When a developer asks about a limit](https://code.claude.com/docs/en/costs#when-a-developer-asks-about-a-limit). For what `/usage` shows a subscriber, see [costs: Using the `/usage` command](https://code.claude.com/docs/en/costs#using-the-/usage-command). For the `rate_limits` field, who receives it and when, see [statusline: Available data](https://code.claude.com/docs/en/statusline#available-data). For the parsed forms, see `check-usage-limit-reset.py` `RESET_RE` (no month token). | 2026-09-29 | That costs section stops covering the reset time; `/usage` stops showing plan usage bars; the statusline page drops `rate_limits` or starts passing it to the model; an in-session account-identity field this skill can read without a sibling plugin ships; or `RESET_RE` starts matching a date-bearing form. |

## Still blocked (limit not yet reset). Hand back, don't busy-wait

While a limit still holds you cannot make progress in this session.
Compose with `/session-flow:handoff` to drop a resume artifact so nothing
is lost, then stop. Automatic wake-and-continue at the reset time is an
external scheduler's job, a desktop scheduled task or a cloud routine
launched to resume from that handoff, not this skill's; keep-going hands
back cleanly and stops. Reach this step only through the live re-check in
the reset bullet above.

## Nothing-off-thread case

If the inventory finds no off-thread work, say so, run step 3's
goal-alignment check when a handoff backs the resume, then go straight to
step 5: reconcile the main thread from its real state and continue. The
interruption may have hit mid-turn on the main thread alone. Recovering
that is still the job.

## What this skill does NOT do

- **Does not diagnose the interruption type.** A short limit, a weekly
  limit, or a crash all take the same recovery, so the cause is not
  classified. (Reading a stated reset time to gate a blocked worker is not
  diagnosis. It does not change the recovery method.)
- **Does not kill or restart live-but-slow work on a hunch.** Action
  follows real output, and kill/restart is gated like any side effect.
- **Does not summarize-and-stall after a limit lifts**. It continues.
- **Does not build or arm its own scheduler.** Still-blocked work is
  handed back by invoking `/session-flow:handoff` via the Skill tool, not parked on a self-armed
  wakeup.
- **Does not trust remembered state.** Every status claim is grounded in
  a fresh read of the real artifact.

## Next

/session-flow:workflow continue

It routes how the recovered session carries on at the next phase boundary.

## Gotchas

- "Probably done" is one failure mode; "it has been a while, kill it" is
  the other. A resumable job that looks finished may have died at 90%; a
  side-effect that looks unsent may have landed just before the cutoff;
  live-but-slow work that looks hung may be one step from done. Read the
  artifact before you decide.
- After a limit lifts, the pull is to summarize and hand back. Resist it:
  if you are running, continue, and report at the end.

Files in this skill

  • SKILL.md11.4 KB
  • evals/evals.json7.3 KB
  • scripts/check-usage-limit-reset.py8.4 KB
  • scripts/check-usage-limit-reset.test.py7.7 KB
  • scripts/check-usage-limit-reset.test.sh212 B
  • scripts/vendor/README.md1 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…