Skip to content
Back to skills

Jsrl Recovery Evaluation

ASecurity

Evaluate Jump-Start Reinforcement Learning recovery evidence with mechanism checks for guide roll-in, curriculum handoff, value update, and proxy metric validity.

  • 247 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 9, 2026
testingpythonbash

Security analysis

A100/100

Scanned September 9, 2026

npx -y skills add VectorSpaceLab/AREX-Skill --skill jsrl_recovery_evaluation --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Jsrl Recovery Evaluation?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Jsrl Recovery Evaluation
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vectorspacelab-jsrl-recovery-evaluation/badge)](https://www.skillsdirectory.com/skills/vectorspacelab-jsrl-recovery-evaluation)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

SKILL.md
---
name: jsrl_recovery_evaluation
description: Evaluate Jump-Start Reinforcement Learning recovery evidence with mechanism checks for guide roll-in, curriculum handoff, value update, and proxy metric validity.
---

# JSRL Recovery Evaluation

## When To Use

Use this skill after a JSRL experiment has run and produced trajectories, training traces, and metrics. It is especially useful in soft-mode reduced recovery, where final scores alone are not enough to prove mechanism faithfulness.

## Inputs

- Module plan target containing dataset, metric, and paper value.
- Recovery metrics for JSRL, vanilla exploration, and optional random switching.
- Trajectory summaries with controller labels.
- Training trace with loss and parameter changes.
- Source-boundary and generated-skill invocation evidence.

## Outputs

- Primary metric such as `success_rate_gain`.
- Mechanism checks for guide use, exploration handoff, curriculum decrease, value update, random ablation, and source-boundary safety.
- Feedback flags for analysis.

## Workflow

1. Confirm recovery metadata matches the module-plan target.
2. Compute JSRL success-rate gain over vanilla exploration.
3. Verify at least one trajectory contains both guide and exploration controllers.
4. Verify curriculum guide steps decrease or the random ablation samples a schedule.
5. Verify the value-update trace changed trainable parameters.
6. Mark reduced recovery explicitly and never label a Q-table proxy as full D4RL/IQL/QT-Opt training.

## Validation

Run:

```bash
python scripts/evaluate_recovery.py --demo
python -m pytest tests
```

The demo emits a compact metric/mechanism bundle.

## Limitations

This skill evaluates evidence; it does not run the environment itself. Recovery harnesses should invoke it after executing the rollout and update skills.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…