Eval-driven skill tuning. Given a task and an LLM-judge rubric, iteratively rewrites a SKILL.md until a downstream executor agent performs well against the judge. Low-code: all evaluation is LLM-as-judge, not deterministic Python.
Installs into .claude/skills of the current project.
Are you the author of Clawpathy Autoresearch?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/stanfish06-clawpathy-autoresearch)
---
name: clawpathy-autoresearch
description: >-
Eval-driven skill tuning. Given a task and an LLM-judge rubric, iteratively
rewrites a SKILL.md until a downstream executor agent performs well against
the judge. Low-code: all evaluation is LLM-as-judge, not deterministic Python.
version: 1.0.0
author: Jay Moore
license: MIT
tags: [meta, autoresearch, skill-tuning, llm-judge, eval-driven]
inputs:
- name: paper_query_or_task
type: string
description: Paper title/URL/PMID/DOI, or a freeform task description
required: true
outputs:
- name: workspace/
type: directory
description: Tuned skill/SKILL.md plus history.jsonl, snapshots, executor_runs
metadata:
openclaw:
requires:
bins: [python3, claude]
always: false
emoji: "π"
homepage: https://github.com/ClawBio/ClawBio
os: [darwin, linux]
trigger_keywords:
- auto research
- autoresearch
- tune a skill
- skill tuning
- improve a skill
- eval-driven
- clawpathy
- replicate paper
- reproduce paper
---
# clawpathy-autoresearch
> [!note] Vault audit 2026-07-24 β USE-10
> Use this for eval-driven tuning that iteratively rewrites an existing SKILL.md against an LLM-judge rubric; to scaffold a new skill from scratch use `skill-builder`, to auto-draft from observed workflows use `autoskill`, to package a plugin bundle use `plugin-creator`. Distinguishing axis: authoring mode (eval-tuning vs manual scaffold vs observation vs plugin packaging).
Eval-driven skill development. The system iteratively rewrites a `SKILL.md`
so a downstream executor agent performs better at a task class, as judged
by an LLM against a paper/task-specific rubric.
## Core idea
```
propose (sonnet) β execute (sonnet, shell) β judge (opus, rubric)
β β
βββββββββ feedback: verdict + recommended edits βββββββββ
```
- **Proposer** rewrites SKILL.md based on the last judge verdict.
- **Executor** runs the new SKILL.md end-to-end inside a workspace.
- **Judge** scores methodology (primary) and outputs (secondary) against
a per-task rubric. Lower is better; 0 = perfect.
- Keep the new SKILL.md only if it strictly beats the best score; else
revert. Stop on target_score or on `early_stop_n` consecutive regressions.
## You are the orchestrator
You (the agent reading this) don't run the loop yourself. You dispatch
subagents to build the workspace, then hand off to the Python loop.
### Phase 1 β Scout
Dispatch a subagent with `prompts/scout.md` to research the paper/task.
Report key findings to the user in a few lines.
### Phase 2 β Scope (you + user)
Have a conversation. Ask ONE question at a time, multiple-choice where
helpful. Agree on:
- what to reproduce / what success looks like
- which data sources are in-bounds
- what methodology expectations belong in the rubric
- iteration budget and target_score (if any)
Present a summary and get approval.
### Phase 3 β Build
Dispatch a builder subagent with `prompts/builder.md` and the agreed
scope. It writes:
- `task.json`
- `rubric.md` β **the authoritative scoring rubric for the LLM judge**
- `reference/` (optional; judge-only)
- `skill/SKILL.md` β seed
Validate:
```python
import sys
from importlib import import_module
from pathlib import Path
# `skills` resolves only from the vault root; the hyphen blocks a plain import.
sys.path.insert(0, str(Path.home() / ".agents"))
validate_workspace = import_module("skills.clawpathy-autoresearch").validate_workspace
print(validate_workspace(Path("WORKSPACE"))) # [] means valid
```
### Phase 4 β Loop
```bash
cd ~/.agents # run from the vault root, else `No module named 'skills'`
python -m skills.clawpathy-autoresearch WORKSPACE_DIR
# or with custom models:
python -m skills.clawpathy-autoresearch WORKSPACE_DIR \
--proposer-model sonnet --executor-model sonnet --judge-model opus
```
The loop streams progress to `WORKSPACE/history.jsonl`, snapshots every
iteration's skill to `WORKSPACE/snapshots/iter-NNN.md`, and writes the
executor's full transcript to `WORKSPACE/executor_runs/iter-NNN.log`.
## Workspace layout
```
workspace/
task.json # task metadata + loop knobs
rubric.md # LLM-judge rubric (the heart of the system)
reference/ # optional ground truth, judge-only
skill/SKILL.md # iterated by the loop
output/ # executor outputs (cleared each iter)
executor_runs/iter-NNN.log # transcripts (judge reads these)
snapshots/iter-NNN.md # per-iter SKILL.md snapshots
history.jsonl # one row per iter: score, kept, verdict
```
## Key principles
- **LLM judge only.** No deterministic Python scorers. All evaluation goes
through `judge.md` + opus. This keeps the system low-code and lets the
rubric carry paper-specific nuance without adding code.
- **Methodology is primary.** The rubric weights "did the agent use sound
methods?" above "did the numbers match?". Ground-truth match is a signal,
not the objective β the goal is better SKILL.md files.
- **Never leak ground truth.** `reference/` is judge-only. The executor
prompt says not to read it, and the judge penalises leakage.
- **No hardcoded answers in SKILL.md.** The proposer prompt and the judge
both enforce this. The executor must derive results by running methods.
- **Snapshots + strict-better revert.** Score on the first iter becomes the
floor. Later iters that tie or regress revert to the best.
## Safety
- All processing is local except scout web fetches for public resources.
- ClawBio disclaimer: research/education tool, not a medical device.
- The executor subagent holds `Bash`, `Write`, and `Edit`, and the instructions it
follows are the proposer's output, not human-reviewed text. It runs once per
iteration, up to `max_iterations` (default 30), and `--parallel K` runs K at once.
- Subagents run under `--permission-mode acceptEdits` by default. `--yolo` switches
them to `bypassPermissions`, removing every approval prompt including for shell
commands. Get the user's explicit consent before passing it, and only for genuinely
unattended runs β the Phase 2 approval covers research scope, not this.
## Gotchas
- **Do not skip scoping.** The rubric is paper-specific; a generic rubric
tunes nothing. Get the user to agree on methodology expectations.
- **Do not write a Python scorer.** Earlier versions of this project did.
They rewarded API-fetching, not methodology. The judge is the scorer.
- **Do not hand-pick the "best" snapshot yourself.** Trust the loop. If
the judge is calibrated wrong, fix the rubric, not the history.