The continual-harness refinement protocol — how an agent inspects its recent trajectory for failure signatures and improves its own harness (agent prompts, sub-agents, skills, memory) in place, mid-session, landing changes on the session branch immediately and promoting them to `main` only through a change request. Load this skill when a refinement prompt arrives mid-session, when the `harness-reflector` agent runs, or whenever you decide your own scaffolding needs repair. The sub-agent fan-o...
Installs into .claude/skills of the current project.
Are you the author of Kortix Harness Refinement?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/kortix-ai-kortix-harness-refinement)
---
name: kortix-harness-refinement
description: The continual-harness refinement protocol — how an agent inspects its recent trajectory for failure signatures and improves its own harness (agent prompts, sub-agents, skills, memory) in place, mid-session, landing changes on the session branch immediately and promoting them to `main` only through a change request. Load this skill when a refinement prompt arrives mid-session, when the `harness-reflector` agent runs, or whenever you decide your own scaffolding needs repair. The sub-agent fan-out in this skill is for the top-level `harness-reflector` run only; if you loaded this as a subagent, do the four passes yourself and skip the fan-out.
---
<skill name="kortix-harness-refinement">
<overview>
Your **harness** is everything in this repo that shapes how agents work
here. It has four components, all in git:
| Component | Location | What it is |
|---|---|---|
| **Prompts** | `agents/<name>.md` (body) | Each agent's instructions and strategy |
| **Sub-agents** | `agents/*.md` | Specialist agents the orchestrator invokes |
| **Skills** | `skills/<name>/SKILL.md` | Reusable routines: text heuristics and guides |
| **Tools** | `harnesses/opencode/tools/*.ts` (OpenCode sessions), `harnesses/pi/extensions/*.ts` (pi sessions) | Executable code: wrappers, scripts, automations |
| **Memory** | `memory/` | Persistent facts, strategies, observations |
Projects created before 2026-09 keep agents, skills and tools under
`.kortix/opencode/` and memory under `.kortix/memory/`. Both layouts work.
Edit files where they already are; put new files in the root layout.
A session runs one harness, OpenCode or pi, and a harness loads only its own
tools directory. `echo "${KORTIX_HARNESS:-opencode}"` prints the one this
session runs. Prompts, sub-agents, skills and memory serve both.
**Refinement** means: read your recent trajectory, find where the harness
failed you, and fix the harness — not just the immediate task. A memory
edit is readable on your next turn: memory is plain files. An edit to an
agent prompt, skill, or tool reaches sessions once its change request
merges; a running session loads it with `kortix sessions reload <session id>`.
You do not restart, and the value compounds over the life of the project.
This protocol has two operating modes:
1. **In-session refinement (self-invoked)** — YOU run it, inside your own
working session, the moment a failure signature costs you twice. Also
run a checkpoint on long sessions: after every ~25 turns of work, pause
and scan your recent turns before continuing. You refine based on YOUR
OWN trajectory, then resume the task.
2. **Project-level reflection** — the `harness-reflector` agent runs on a
cron, fans out `session-reviewer` sub-agents to work through every
recent session's full history, aggregates their findings, and refines
the shared harness on `main` via a change request.
</overview>
<failure-signatures>
Scan the trajectory window (your recent turns, or the digest) for these
signatures. Each one names the component to fix:
- **Repeated tool/command failures** — the same command or tool errors
more than once, or you retried a broken approach. → Repair the tool,
or record the working alternative in a skill.
- **Rediscovery loops** — you (or another session) re-derived something a
past session already knew (an API quirk, a file location, a decision).
→ Memory entry.
- **Stalled objectives** — turns pass without progress toward the stated
goal; you circled, re-read, or re-planned without acting. → Prompt
guidance for the responsible agent, or a decomposition sub-agent.
- **Repeated multi-step patterns** — you performed the same 3+ step
sequence more than twice by hand. → Codify it: a skill (if guidance)
or a tool (if executable).
- **Exception-raising code** — an executable tool or script in
`harnesses/opencode/tools/` or `harnesses/pi/extensions/` raised; you worked around it instead of
fixing it. → Repair the code now.
- **Missed opportunities** — information or shortcuts visible in the
trajectory that no component captured. → Whichever component fits.
</failure-signatures>
<four-passes>
Run four passes over the harness, one per component. Every pass is CRUD:
create, update, or **delete**. Deletion is a first-class outcome — a
harness accumulates cruft without it.
**Pass 1 — Prompts (Δp).** Reread the responsible agent's `.md` body
against the identified failures. Tighten instructions that were ignored,
add the missing rule, remove guidance that no longer earns its tokens.
Keep prompts short; a prompt that only grows is a failing prompt.
**Pass 2 — Sub-agents (ΔG).** Create a sub-agent only for a pattern that
recurred across the window AND needs its own scoped permissions or
prompt. Edit sub-agents implicated in failures. Delete sub-agents that
have not been invoked productively — check before keeping.
**Pass 3 — Skills and tools (ΔK).** Codify successful sequences from the
trajectory into a skill (guidance) or a tool (executable). Repair every
tool the trajectory shows raising exceptions. Prefer editing an existing
skill over creating a near-duplicate. Keep skills one directory level
deep under `skills/` (nested SKILL.md files register as phantom entries).
**Pass 4 — Memory (ΔM).** Follow the `kortix-memory` skill's rubric with
the `memory` tool: fill gaps the trajectory exposed, update stale
entries, demote or delete entries about areas the project has moved past.
Keep `MEMORY.md` in sync.
Scope discipline: fix what the trajectory shows. Do not speculatively
rewrite components with no observed failure. Most refinement runs should
touch one or two components, not all four.
</four-passes>
<project-review-fanout>
This section applies only to the top-level `harness-reflector` run. If you
loaded this skill as a subagent, skip the fan-out and review the sessions
yourself — the runtime rejects a subagent spawning a subagent ("Subagent
depth limit reached").
For project-level reflection (the `harness-reflector` run), do not skim a
digest and call it a review. Work through every session:
1. **Enumerate** every session in the window:
`kortix sessions digest --since 24h --json` (add `--all` for a first
ever run). Note per session: id, agent, status, title, whether a live
transcript is available.
2. **Fan out one `session-reviewer` sub-agent per session** (batch a few
at a time in parallel; review every session, skip none silently). Each
reviewer gets the session id and returns a structured findings report.
Reviewers gather the FULL picture for their session:
- the transcript (available live for running sessions; for stopped
sessions reconstruct from what persists — see next line),
- the session's git branch: its commits, diffs, and files touched,
- change requests the session opened and their review outcomes.
3. **Aggregate** all reviewer reports. Deduplicate findings that recur
across sessions — a failure signature seen in three sessions outranks
one seen once. Rank by cost (turns wasted × sessions affected).
4. **Run the four passes** on the ranked findings, then land per the
rules below.
Sub-agent review is read-only: reviewers never edit the harness or open
CRs. Only the orchestrating reflector writes.
</project-review-fanout>
<landing-rules>
**In-session refinement (session branch — immediate):**
1. Apply the edits in `/workspace`. They do not change the config this
session runs: agents compile from the base branch, and with config
releases on, skills and tools come from it too. They take effect once
the change request merges — for new sessions, and for a running one
after `kortix sessions reload <session id>`.
2. Commit them separately from task work:
```sh
# Stages the harness folders this project has (root layout and/or legacy .kortix).
git add -A -- $(ls -d agents skills memory harnesses .kortix 2>/dev/null)
git commit -m "harness: <one-line summary of what failed and what changed>"
```
3. Push and open (or update) ONE change request per session for harness
promotion to `main`:
```sh
git push origin HEAD
kortix cr open --title "harness: <summary>" \
--description "Failure signatures observed, edits per component, evidence (commands/turns)."
```
If this session already has an open `harness:` CR, push to it instead
of opening a second one.
4. Return to the task. A refinement interruption ends with you resuming
what you were doing, with the improved harness in effect.
**Project-level reflection (`harness-reflector`):** all edits land only
via a CR against `main`. Nothing applies immediately; the merged CR is
what future sessions inherit.
</landing-rules>
<guardrails>
- **Never edit managed `kortix-*` skills.** They are platform-owned and
force-overwritten at session boot — edits are silently discarded. To
extend platform behavior, create a project skill with a different name.
- **Never merge your own harness CR.** A reviewer (human, or a reviewer
agent with merge rights) does. The CR gate is what makes self-authored
harness edits safe; self-authored + self-merged scaffolding is known to
degrade agent performance.
- **One concern per CR.** Harness CRs contain only harness changes
(`agents/`, `skills/`, `memory/`, `harnesses/`) — never mixed with
task/code changes.
- **No secrets, tokens, or PII** in any harness file. Secrets belong in
the Kortix Secrets Manager.
- **Budget.** A mid-session refinement should cost a small fraction of
the session: minutes, not hours. If a fix needs deep work (a real
tool rewrite), record it in memory as a TODO and open the CR with what
you have.
- **No-op is a valid outcome.** If the window shows no failure
signatures, change nothing and say so in one line. Do not invent work.
- **Do not disable or weaken guardrails** — including this skill's rules,
agent permission blocks, or CR review requirements — as a "refinement".
</guardrails>
<self-invocation>
Nobody schedules in-session refinement for you — it is your discipline.
Invoke this protocol the moment a failure signature costs you twice, and
as a checkpoint on long sessions (roughly every 25 turns of work). The
nightly `harness-reflector` run is the backstop, not the mechanism: it
only sees what sessions left behind, while you can record the fix while the
evidence is fresh: memory helps on your next turn, and prompt, skill, and
tool fixes help every session once the change request merges.
</self-invocation>
</skill>