Skip to content
Back to skills

Clawpathy Autoresearch

ASecurity

Eval-driven skill tuning. Given a task and an LLM-judge rubric, iteratively rewrites a SKILL.md until a downstream executor agent performs well against the judge. Low-code: all evaluation is LLM-as-judge, not deterministic Python.

  • 8 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 12, 2026
researchpythonrustgoshellbashgitapi

Works with

  • api

Security analysis

A100/100

Pro scans all 19 files and shows the line behind each finding

Scanned September 12, 2026

npx -y skills add stanfish06/skillquarium --skill clawpathy-autoresearch --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Clawpathy Autoresearch?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Clawpathy Autoresearch
[![Security: A β€” Skills Directory](https://www.skillsdirectory.com/api/skills/stanfish06-clawpathy-autoresearch/badge)](https://www.skillsdirectory.com/skills/stanfish06-clawpathy-autoresearch)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: clawpathy-autoresearch
description: >-
  Eval-driven skill tuning. Given a task and an LLM-judge rubric, iteratively
  rewrites a SKILL.md until a downstream executor agent performs well against
  the judge. Low-code: all evaluation is LLM-as-judge, not deterministic Python.
version: 1.0.0
author: Jay Moore
license: MIT
tags: [meta, autoresearch, skill-tuning, llm-judge, eval-driven]

inputs:
  - name: paper_query_or_task
    type: string
    description: Paper title/URL/PMID/DOI, or a freeform task description
    required: true

outputs:
  - name: workspace/
    type: directory
    description: Tuned skill/SKILL.md plus history.jsonl, snapshots, executor_runs

metadata:
  openclaw:
    requires:
      bins: [python3, claude]
    always: false
    emoji: "πŸ”"
    homepage: https://github.com/ClawBio/ClawBio
    os: [darwin, linux]
    trigger_keywords:
      - auto research
      - autoresearch
      - tune a skill
      - skill tuning
      - improve a skill
      - eval-driven
      - clawpathy
      - replicate paper
      - reproduce paper
---

# clawpathy-autoresearch

> [!note] Vault audit 2026-07-24 β€” USE-10
> Use this for eval-driven tuning that iteratively rewrites an existing SKILL.md against an LLM-judge rubric; to scaffold a new skill from scratch use `skill-builder`, to auto-draft from observed workflows use `autoskill`, to package a plugin bundle use `plugin-creator`. Distinguishing axis: authoring mode (eval-tuning vs manual scaffold vs observation vs plugin packaging).

Eval-driven skill development. The system iteratively rewrites a `SKILL.md`
so a downstream executor agent performs better at a task class, as judged
by an LLM against a paper/task-specific rubric.

## Core idea

```
  propose (sonnet)  β†’  execute (sonnet, shell)  β†’  judge (opus, rubric)
       ↑                                                       β”‚
       └──────── feedback: verdict + recommended edits β”€β”€β”€β”€β”€β”€β”€β”€β”˜
```

- **Proposer** rewrites SKILL.md based on the last judge verdict.
- **Executor** runs the new SKILL.md end-to-end inside a workspace.
- **Judge** scores methodology (primary) and outputs (secondary) against
  a per-task rubric. Lower is better; 0 = perfect.
- Keep the new SKILL.md only if it strictly beats the best score; else
  revert. Stop on target_score or on `early_stop_n` consecutive regressions.

## You are the orchestrator

You (the agent reading this) don't run the loop yourself. You dispatch
subagents to build the workspace, then hand off to the Python loop.

### Phase 1 β€” Scout

Dispatch a subagent with `prompts/scout.md` to research the paper/task.
Report key findings to the user in a few lines.

### Phase 2 β€” Scope (you + user)

Have a conversation. Ask ONE question at a time, multiple-choice where
helpful. Agree on:
- what to reproduce / what success looks like
- which data sources are in-bounds
- what methodology expectations belong in the rubric
- iteration budget and target_score (if any)

Present a summary and get approval.

### Phase 3 β€” Build

Dispatch a builder subagent with `prompts/builder.md` and the agreed
scope. It writes:
- `task.json`
- `rubric.md` β€” **the authoritative scoring rubric for the LLM judge**
- `reference/` (optional; judge-only)
- `skill/SKILL.md` β€” seed

Validate:
```python
import sys
from importlib import import_module
from pathlib import Path

# `skills` resolves only from the vault root; the hyphen blocks a plain import.
sys.path.insert(0, str(Path.home() / ".agents"))
validate_workspace = import_module("skills.clawpathy-autoresearch").validate_workspace
print(validate_workspace(Path("WORKSPACE")))  # [] means valid
```

### Phase 4 β€” Loop

```bash
cd ~/.agents   # run from the vault root, else `No module named 'skills'`
python -m skills.clawpathy-autoresearch WORKSPACE_DIR
# or with custom models:
python -m skills.clawpathy-autoresearch WORKSPACE_DIR \
  --proposer-model sonnet --executor-model sonnet --judge-model opus
```

The loop streams progress to `WORKSPACE/history.jsonl`, snapshots every
iteration's skill to `WORKSPACE/snapshots/iter-NNN.md`, and writes the
executor's full transcript to `WORKSPACE/executor_runs/iter-NNN.log`.

## Workspace layout

```
workspace/
  task.json                  # task metadata + loop knobs
  rubric.md                  # LLM-judge rubric (the heart of the system)
  reference/                 # optional ground truth, judge-only
  skill/SKILL.md             # iterated by the loop
  output/                    # executor outputs (cleared each iter)
  executor_runs/iter-NNN.log # transcripts (judge reads these)
  snapshots/iter-NNN.md      # per-iter SKILL.md snapshots
  history.jsonl              # one row per iter: score, kept, verdict
```

## Key principles

- **LLM judge only.** No deterministic Python scorers. All evaluation goes
  through `judge.md` + opus. This keeps the system low-code and lets the
  rubric carry paper-specific nuance without adding code.
- **Methodology is primary.** The rubric weights "did the agent use sound
  methods?" above "did the numbers match?". Ground-truth match is a signal,
  not the objective β€” the goal is better SKILL.md files.
- **Never leak ground truth.** `reference/` is judge-only. The executor
  prompt says not to read it, and the judge penalises leakage.
- **No hardcoded answers in SKILL.md.** The proposer prompt and the judge
  both enforce this. The executor must derive results by running methods.
- **Snapshots + strict-better revert.** Score on the first iter becomes the
  floor. Later iters that tie or regress revert to the best.

## Safety

- All processing is local except scout web fetches for public resources.
- ClawBio disclaimer: research/education tool, not a medical device.
- The executor subagent holds `Bash`, `Write`, and `Edit`, and the instructions it
  follows are the proposer's output, not human-reviewed text. It runs once per
  iteration, up to `max_iterations` (default 30), and `--parallel K` runs K at once.
- Subagents run under `--permission-mode acceptEdits` by default. `--yolo` switches
  them to `bypassPermissions`, removing every approval prompt including for shell
  commands. Get the user's explicit consent before passing it, and only for genuinely
  unattended runs β€” the Phase 2 approval covers research scope, not this.

## Gotchas

- **Do not skip scoping.** The rubric is paper-specific; a generic rubric
  tunes nothing. Get the user to agree on methodology expectations.
- **Do not write a Python scorer.** Earlier versions of this project did.
  They rewarded API-fetching, not methodology. The judge is the scorer.
- **Do not hand-pick the "best" snapshot yourself.** Trust the loop. If
  the judge is calibrated wrong, fix the rubric, not the history.

Files in this skill

  • .gitignore957 B
  • SKILL.md6.6 KB
  • __init__.py383 B
  • __main__.py1.8 KB
  • dispatcher.py2.7 KB
  • examples/champions/README.md1.4 KB
  • examples/champions/trubetskoy_scz_finemap_0.235.md49.8 KB
  • examples/champions/yengo_height_ldsc_h2.md25.7 KB
  • examples/demo_task/README.md690 B
  • examples/demo_task/rubric.md1.7 KB
  • examples/demo_task/skill/SKILL.md331 B
  • examples/demo_task/task.json463 B
  • examples/presentation_intro.md3.6 KB
  • executor.py1.1 KB
  • judge.py2.6 KB
  • loop.py5.1 KB
  • loop_parallel.py9.2 KB
  • multitask_loop.py15.1 KB
  • prompts/builder.md2 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…