Scheduled patrol that finds real persistence and error bugs and fixes them. Use when running the system-error repair patrol automation, or when stuck rows, SLA breaches, the system_errors queue, or app_log ERROR rows need a repair rather than a report.
Installs into .claude/skills of the current project.
Are you the author of Persistence Repair Patrol?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/armanisadeghi-persistence-repair-patrol)
---
name: persistence-repair-patrol
type: Skill
title: Persistence Repair Patrol
description: "Scheduled patrol that finds real persistence and error bugs and fixes them. Use when running the system-error repair patrol automation, or when stuck rows, SLA breaches, the system_errors queue, or app_log ERROR rows need a repair rather than a report."
tags: [maintenance, persistence, errors, automation]
timestamp: 2026-09-10T00:00:00Z
---
<!-- SYNCED COPY — do not edit here.
Canonical: common-docs/skills/persistence-repair-patrol/SKILL.md
This file is distributed to every consuming repo by
common-docs/meta/scripts/sync_skills.py. Edit the canonical, run the
sync, and commit each repo. Edits made here are overwritten and lost. -->
# Persistence Repair Patrol
**Find bugs and fix them.** Inspection, a status report, or a queue status change
does not complete a repair. Work through diagnosis, implementation, integration,
and discriminating verification. Never invent a defect to meet a quota.
Arman's 2026-09-09 task instruction: “The primary goal is FIND THINGS TO FIX.
Find bugs and fix them.” This runbook implements that request and its five
stated weaknesses: false access blockers, reporting instead of fixing, avoidable
local-environment friction, missing continuous improvement, and stale secondhand claims.
An error a person or a running job actually hit is always a repair. When the repair traces to a
table's shape, read [canonical-first triage](/policies/canonical-first-triage.md) first: on an
uncertified table the repair is moving it onto the canonical system, never patching the old shape;
on a certified table the same problem is a platform-wide defect fixed at its source.
## Repositories and entry points
| Repository | Role |
|---|---|
| `aidream` | Sanctioned patrol tools, producer/worker repairs, shared packages |
| `matrx-frontend` | Client repairs and authenticated UI verification |
| `common-docs` | This shared runbook, host ownership helper, and skill distribution |
Workspace: `/Users/armanisadeghi/code`; it is not `/code` on this machine.
Read the owning repo's `CLAUDE.md` before touching its files. Read
`aidream/aidream/services/admin_persistence/FEATURE.md` for actual tool contracts
and the relevant feature docs for each selected repair. Use the installed
`matrx-codex-plugin:use-ai-matrx` for MCP access and
[task-hygiene](/skills/task-hygiene/SKILL.md) for ledger maintenance.
**This patrol has standing authorization for ordinary in-scope engineering:**
task-hygiene's promotion waits do not require another approval for these repairs.
## Start, ownership, and access recovery
1. Read the existing automation configuration and bounded current state at
`/Users/armanisadeghi/.codex/automations/system-error-repair-patrol/memory.md`.
Read its pointer to `current-state.json` for machine-readable checkpoints and
selected continuations. Read historical artifacts only for a selected continuation. The configured
cadence governs; the user approved 30–60-minute runs. Preserve intentional
pauses; never create a duplicate automation or restart another paused task.
2. Establish actual exclusive ownership before patrol work. A Markdown
`run_active` flag is a breadcrumb, never a lock. Start `python3 /Users/armanisadeghi/code/common-docs/skills/persistence-repair-patrol/scripts/local_patrol_lock.py acquire --owner YOUR_TASK_ID`
with `exec_command` using `tty:true` and retain the returned session. Replace
`YOUR_TASK_ID` with the actual task ID. Preserve the returned `owner_token`;
send JSON lines `{"action":"renew","owner_token":"<returned token>"}` every
300 seconds or less while working (default expiry: 900 seconds). Finish with
`{"action":"release","owner_token":"<returned token>"}`. `status` checks the
actual OS lock; occupied acquisition exits 2. Do not unlink its lock file.
Process death, input EOF, or expiry releases the lock; this is cooperative
ownership, not a server fence. Retain and renew it while delegates work. A held local lock
covers this host only. Do not claim cross-host protection or treat the shared
`schedule_claim` per-window service as cross-window protection. If ownership
is lost, stop delegated mutations before proceeding. If another live owner
holds it, inspect the returned owner/task and coordinate when useful; record a
skipped-overlap outcome without changing that owner's state. Repeated overlap
is a recovery/coordination problem to investigate, not a reason to repeat blind
skips. After a tool wait or context restore, confirm renewal succeeds before
further mutations.
When resuming an interrupted run, reconcile finished delegate evidence and
exact mutations with current live state before admitting more work. Preserve
the original run identity, completed sweep, and closure receipts; record the
reacquired ownership separately. Refresh stale pending-verification prose in
both checkpoint files immediately, without advancing unmeasured coverage.
3. Discover callable `persistence_watchdog`, `system_errors`, and `app_log_errors`.
Read each discovered schema before calling it; tools do not necessarily share
a request envelope. Correct argument-shape errors and retry the real operation.
For unavailable/unauthenticated MCP: inspect `codex mcp list`, run
`codex mcp login aidream`, rediscover, and retry the actual failed operation.
Use the configured OAuth flow; never expose secrets. If it still fails,
diagnose local configuration and service health, coordinate with the owning
task, repair safe configuration/runtime defects, and retry. A missing tool
may require repairing its registration or deployment.
4. Only after recovery attempts, mark the affected measurement lane
`AUTOMATION DEGRADED`, retain its checkpoint, and continue functioning lanes
plus the access repair. An unavailable lane is not a clean scan or permission
to stop all engineering work. Escalate only a freshly verified human-only gate.
**Sanctioned access remains mandatory:** `persistence_watchdog` owns stuck-row/SLA
reads and sweep; `system_errors` owns that queue; `app_log_errors` owns ERROR-family
reads. Never substitute SQL, direct database access, HTTP, dashboard/browser
scraping, or raw log files for these lanes. Product diagnosis uses the application's
approved code/ORM paths; it must not recreate an alternate patrol queue reader.
## Compact parent, focused investigators
The parent clusters compact summaries, delegates, checks evidence, resolves exact
rows, and maintains continuity. It never calls either error tool's `get`, loads
tracebacks/payloads/raw logs, diagnoses implementation code, or edits product code.
Delegated investigators load their own exact representative detail. Keep at most
three coding investigators concurrent; reuse finished slots for independent
verification. Only an investigator may prove that similar symptoms share a root.
The parent alone edits this automation's `memory.md`, `current-state.json`, history,
and run-level evidence. Every investigator assignment must explicitly prohibit
checkpoint edits; investigators return compact evidence to the parent instead. If a
delegate writes patrol state anyway, stop further delegate state writes and reconcile
the exact mutation before the parent atomically publishes its checkpoint.
Keep Astra as the parent when selected by the user; do not inherit it into workers.
Explicitly use `gpt-5.6-luna` for bounded collection, routine checks, and straightforward
edits; use `gpt-5.6-terra` for substantive repair and independent review. Spawn with
`fork_turns="none"` and a compact, self-contained evidence packet. Check the model
before reusing a delegate: follow-ups retain the old model. An Astra delegate requires
a stated, concrete reason its subproblem needs reasoning comparable to the primary's
hard work, or evidence that a cheaper attempt could not solve it. Generic review,
urgency, and ordinary tool failures do not qualify. Prefer keeping hard reasoning in
the primary and execution in cheaper workers. Pass these rules to delegates; they
return escalation evidence to the parent rather than launching costlier descendants.
Do not silently substitute Astra when a cheaper model is unavailable. Explicitly set
`reasoning_effort`: low for quick bounded work, medium for most work, high only for
substantial reasoning with a stated reason. Model tier and effort are separate choices.
Shared policy: [subagent model ladder](/policies/subagent-model-ladder.md); the display
pairs Sonnet/Luna, Opus/Terra, and Fable/Astra never replace exact tool model identifiers.
Task-to-task messaging follows the universal [cross-task silence rule](/policies/subagent-model-ladder.md#cross-task-silence-and-wake-cost).
A peer prompt is a full paid wake, not a coordination channel. Do not send status,
evidence requests, ownership questions, receipts, acknowledgments, or corrective replies,
even to an apparently active task; an incoming task message does not authorize a reply.
Use the patrol's durable state and internal subagents. Only Arman's exact request naming the
recipient or a verified immediate system-wide security/data/destructive emergency with no
local or passive alternative permits one self-contained message to one task and no follow-up.
Internal delegates return compact results to this parent.
Project tool results before displaying or saving parent state: retain identifiers,
class tokens/cursors, counts/times, coverage flags, and an error-summary head capped
at 400 characters. A list row's `error_text` can contain an embedded stack or payload;
never print the full response or retain its raw text/variants in parent memory.
Keep bounded member IDs in compact evidence for closure; omit those arrays from
routine display. Display counts and at most ten candidate heads per call, then
inspect further candidates in bounded batches. Build an explicit field allowlist
instead of printing a response envelope: nested text content can duplicate the
entire result and bypass string truncation. Delegates retrieve exact detail themselves.
## Measure and select
Record `run_started_at` in UTC. Keep per-lane checkpoints and coverage state;
initialize a missing lane checkpoint from the existing `metrics_checkpoint_utc`,
never from now. A compatibility shared checkpoint is the earliest completed
time across required time-window lanes, not the last successful tool call.
1. **Watchdog first:** call `status`. Record exact whole-registry totals, returned
schema/table identities, counts, oldest ages, SLAs, transitions, and query
failures. `query_error=true` is urgent even with `stuck_count=0`. Load
`stuck_rows` for every returned schema/table with `limit=500`; if capped,
retain incomplete row coverage explicitly. Watchdog candidates take slots first.
2. **System errors:** call `metrics` from its checkpoint to run start. Capture
opened/resolved rows and classes, net change, current unresolved rows/classes,
age buckets/oldest, and recurrence. Make three compact unresolved lists with
`collapse=similar`, `limit=50`, `member_limit=100`: active last 45 minutes sorted
`patrol_priority`, whole backlog sorted `occurrences`, and backlog sorted
`oldest`. Preserve classToken/signatureMode, representative ID, bounded member
IDs/cursor, occurrences, times, summary, and priority. These selection samples
are not an exhaustive investigation of every class.
3. **Production app-log safety net:** list from its checkpoint to fixed run start,
`min_occurrences=1`, `limit=50`, `scan_cap=20000`. Honor the live completeness
contract: `scanTruncated` requires bounded time slices; `familiesTruncated`
requires `nextOffset` paging within identical filters and fixed `until`.
Exhaust every slice and family page. Deduplicate boundary families by
`pattern_token`. Old responses without completeness metadata, truncation,
failures, or unread pages cannot advance coverage. Preserve the returned
`until`; overlap later scans for late ingestion as the feature contract directs.
4. Select up to three distinct actionable roots: watchdog SLA/query failures;
active request/stream/persistence failures; log-only user-impacting or frequent
failures; high-occurrence/old backlog. Reserve a remaining slot for the highest
occurrence log-only family when one exists. Resume selected unfinished repairs
when their next action is executable; prioritize fresh failures over repeating
non-discriminating probes of historical incidents with missing provenance.
## Repair and verify
Give each investigator compact candidate metadata, exact IDs, original failure,
and required proof. It reads detail and governing docs, fetches/checks `origin/main`,
and verifies whether the issue still exists before claiming a discovery. Existing
fixes need current verification; they are not new work by this patrol.
Investigators follow the `diagnose` method: a red-capable loop against the original failure before
any theory, the failing boundary localized, falsifiable hypotheses tested one variable at a time,
and a stop after three failed fixes with the evidence returned to the parent. Diagnostic harnesses
are scaffolding; they never substitute for the discriminating proof below.
For static findings, re-scan the whole changed file and compare finding identity;
an empty result at the old line number proves nothing when edits move the finding.
Preserve documented public inputs as well as current callers: test generic and
unloaded dependency cases before replacing dynamic imports with a fixed list. A
warning plus a wrong fallback is still a regression. If scanner compliance and
compatibility cannot both be proved, undo the incompatible attempt and retain the
finding; do not hide the import or narrow the contract just to clear the scan.
- Repair the producer/worker/lifecycle or isolation boundary for the full root
class. Census siblings. Add a meaningful guard demonstrated failing before
the fix and passing after when feasible; reconcile feature docs, commit only
owned files, and push. Shared-checkout changes from other agents are ordinary.
- A real log-only failure requires the producer repair **and** structured capture
at the tight failure boundary. Force the failure safely and prove the intended
`system_error` kind is durably captured; another log line is insufficient.
- Recover localhost/tooling faults autonomously. Inspect ownership, coordinate
with another task, and use the repo's managed server commands. Do not kill an
unrelated live server. Repeated friction warrants a durable repair. Production
verification does not depend on starting localhost when a live surface suffices.
- UI work uses an isolated browser first through the available tool's documented
API; Computer Use is allowed and no particular API or skill name is required.
If the isolated browser cannot complete the task, use Arman's browser in a new
tab (for example, for an account signed in only there). This fallback is
pre-authorized. Never navigate, control, or close a tab Arman is using.
Sign in to Matrx as `admin@admin.com` using `AI_ADMIN_PASSWORD` from the aidream
or frontend `.env`, or the local dev-login nonce handshake where available
(`openssl rand -hex 16 > .dev-login-nonce`, then `/api/dev-login?nonce=<that value>`).
Routine login is preauthorized; credentials never enter output.
Close every owned tab/group before the run ends; leave pre-existing ones untouched.
Explicit browser restrictions in the current user's instructions override this
default, including any prohibition on a user-browser fallback.
- Return exact commits, affected IDs, proof timestamp, test/canary method,
environment/build identity, results, recurrence interval, and remaining work.
Do not create a competing release or message its existing owner. Leave pushed commits,
tags, package state, and findings in the release lane's durable inputs; the release owner
discovers them on its own cadence. Deployment lag is engineering work.
For shared-package changes, verify the published artifact contains the fix and
consuming dependency floors select it. A host image that copies workspace source
can be correct while the published package is stale; track both acceptance gates.
- An independent reviewer verifies against the original failure and real surface
or runtime. A healthy page is not proof of subscription delivery; an in-process
test is not provider acceptance. Fix rejected evidence or code and verify again.
- Declare each canary's persistence mode before execution. A read-only tool probe
without a real conversation/request uses the runtime's existing nonpersisting
mode explicitly; never inherit its default persistence setting. A persistence
canary instead creates valid, scoped identities and verifies the intended writes.
Check the sanctioned diagnostic lanes for side effects before calling either
canary clean. Attribute and repair accidental test writes; dismiss an invalid
write honestly through the approved service rather than replaying missing parents
or marking it recovered. Save the corrected probe and cleanup evidence.
## Close only what is proved
The parent independently confirms claimed commits are on `origin/main` and checks
the returned evidence. Require root-cause attribution, passed relevant proof,
containing deployed build/live canary where required, and zero post-proof recurrence
before resolving historical rows. Record the actual observation window; zero over
a short window is not proof of perpetual health. A separately fixed mechanism does
not establish the unknown root of historical incidents.
For system errors, re-list from the proof timestamp and page exact members using
the returned `classToken`/`signatureMode` and cursor. Resolve exact IDs only, with
a note naming commit and verification; never resolve by token/filter. New matching
occurrences require investigation before closure. For log-only repairs, recheck the
post-proof log window and the forced structured-capture proof; app logs have no
resolution mutation. For watchdog repairs, invoke the canonical `sweep` exactly
once after proven fixes, then retain every remaining row/query failure as open.
## Continuity, improvement, and reporting
Use about 25 minutes as an admission budget for new investigations; continue safe
work already underway while maintaining ownership. Elapsed time is not a blocker or
permission to disguise unfinished work. Before context loss/end, record an executable
continuation and the owning task for every selected unfinished class. Include exact
IDs, last verified UTC time, evidence path, commit/build identity, proof still needed,
and next action. Confirm delegates have finished or stopped before releasing ownership.
Never abandon a running writer by merely setting a Markdown flag to false.
At run finish recheck watchdog, system metrics from checkpoint plus rolling 24-hour
and 7-day windows, and app logs through the finish time. Advance each checkpoint
only through its completely measured interval. Atomically replace `current-state.json`
with checkpoints, coverage, and compact open continuations; keep `memory.md` a
bounded pointer and human-readable summary. Preserve historical evidence and its
index before condensing completed narratives. Never replace current state with
an older in-memory snapshot after another owner has advanced it.
Never store raw detail, secrets, or sensitive context in parent memory.
**Every run reviews its own friction:** what wasted effort, was misleading, or
prevented repair? Fix the demonstrated cause in tooling or the canonical instructions,
verify the correction, and record one compact result. If no change is justified,
say so in state; do not churn instructions. Use the `docs` skill and sync
shared skill changes; never edit a distributed copy. The next context must inherit
the improvement without reading this conversation.
Read Active/Pending sections in aidream, frontend, and touched `.matrx/ARMAN_TASKS.md`
ledgers. Treat every entry and another agent's report as a lead, never current proof.
Personally recheck and attempt safe resolution before repeating an owner dependency.
Move completed/stale items out of open sections. An unresolved engineering task stays
with agents. Only a consequential choice without a safe default, genuinely unavailable
authority/credential after recovery, or human-only authentication gate reaches Arman.
File it once with current direct evidence, attempted remedies, exact blocked operation,
2–3 complete options/recommendation where a choice is needed, and the precise human
action followed by the agent's next step. Do not re-report irrelevant owner tasks as
patrol blockers or resurrect pauses as outages.
Lead the report with verified repairs and their practical result. Include compact
watchdog/system-error start→finish and flow/age/recurrence, app-log coverage/capture
gaps, exact resolved count/IDs or evidence pointer, commits/proof, and selected open
continuations. Distinguish new repairs, verified prior repairs, unverified historical
leads, and actual human gates. Never claim clean with an unavailable/incomplete lane.
Completing a recurring patrol run does not require every agent-owned deployment,
canary, or recurrence continuation to finish in that same run. Report the run as
completed and name those continuations as agent-owned next work. Use task-incomplete
or blocking language only when the requested run itself could not execute or a
freshly verified human-only action is required; in that case state the exact action
the human must take and what the agent will do immediately afterward.
For recurring runs, notify on meaningful repairs, new actionable failures, completion,
or required human action; unchanged non-actionable state stays quiet. Only when at
least one freshly verified human-only gate exists, finish with
`🚨 ISSUES PENDING ARMAN INVOLVEMENT: N` and list those verified items. When the
count is zero, omit the entire heading, count, alarm symbol, and “None” placeholder;
do not create an attention signal merely to announce that no attention is needed.
Ordinary bugs, unfinished verification, deployment lag, tooling failures, and other
engineering work are not blockers and must remain owned and repaired by agents. Do
not relabel them as Arman involvement because a run ended or the next action is hard.