Skip to content
Back to skills

Ai Coscientist

ASecurity

Train LLMs to generate high-quality research plans via rubric-based RL without requiring experimental verification. Extracts research goals and domain-specific rubrics from scientific papers, uses frozen model as grader with 12-22% relative improvements, achieves human-expert preference 70% of time with strong cross-domain generalization.

  • 6 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 9, 2026
researchpythongo

Security analysis

A100/100

Scanned September 9, 2026

npx -y skills add ADu2021/skillXiv --skill ai-coscientist --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Ai Coscientist?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Ai Coscientist
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/adu2021-ai-coscientist/badge)](https://www.skillsdirectory.com/skills/adu2021-ai-coscientist)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: ai-coscientist
title: "Training AI Co-Scientists Using Rubric Rewards"
version: 0.0.2
engine: skillxiv-v0.0.2-claude-opus-4.6
license: MIT
url: https://arxiv.org/abs/2512.23707
keywords: [research, reinforcement-learning, rubric-rewards, cross-domain]
description: "Train LLMs to generate high-quality research plans via rubric-based RL without requiring experimental verification. Extracts research goals and domain-specific rubrics from scientific papers, uses frozen model as grader with 12-22% relative improvements, achieves human-expert preference 70% of time with strong cross-domain generalization."
---

## Overview

Automated research plan generation using extractable domain knowledge from papers.

## Core Technique

**Rubric Extraction and Grading:**

```python
# Extract from papers
research_rubrics = extract_rubrics_from_papers(papers)

# Grade via frozen model
grade = frozen_model.score(generated_plan, rubrics)

# GRPO training on grade signal
```

## When to Use

Use when: Research automation, domain-specific planning, cross-domain generalization.

## References

- Rubric extraction from scientific papers
- Frozen model as grader
- Self-reward GRPO training

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…