Skip to content
Back to skills

Pql Reduced Recovery Harness

ASecurity

Run a mechanism-faithful reduced Parallel Q-Learning proxy experiment with replay, mixed exploration, and optimizer evidence.

  • 247 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 9, 2026
testingpythonbash

Security analysis

A100/100

Scanned September 9, 2026

npx -y skills add VectorSpaceLab/AREX-Skill --skill pql_reduced_recovery_harness --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Pql Reduced Recovery Harness?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Pql Reduced Recovery Harness
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vectorspacelab-pql-reduced-recovery-harness/badge)](https://www.skillsdirectory.com/skills/vectorspacelab-pql-reduced-recovery-harness)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

SKILL.md
---
name: pql_reduced_recovery_harness
description: Run a mechanism-faithful reduced Parallel Q-Learning proxy experiment with replay, mixed exploration, and optimizer evidence.
---

# PQL Reduced Recovery Harness

Use this skill when full Isaac Gym reproduction is blocked but soft-mode recovery allows a declared reduced proxy. The run must still exercise PQL's core mechanism: many actors, mixed exploration, replay sampling, separate value and policy updates, and numeric training evidence. Do not use it to claim full benchmark reproduction.

## Inputs

- Paths to generated PQL topology, mixed exploration, and diagnostics skills.
- Actor count, rollout length, replay capacity, learning rates, and random seed.
- Module-plan target metadata for populating recovery results.

## Outputs

- Training trace with `loss_before`, `loss_after`, `params_before`, and `params_after`.
- Mechanism checks for actor coverage, mixed exploration, replay use, and optimizer execution.
- Numeric proxy metrics such as `loss_reduction` and final return estimate.

## Workflow

1. Import the topology, mixed exploration, and diagnostics helpers.
2. Build a scalar-control task where reward is highest near a target action.
3. Use many actors with round-robin mixed exploration to collect replay transitions.
4. Fit a simple critic parameter toward Bellman-style rewards.
5. Update a policy parameter toward the critic-implied target.
6. Record every parameter change and emit recovery-compatible JSON artifacts.

## Validation

Run:

```bash
python scripts/reduced_pql.py --actor-count 32 --rollout-steps 4 --updates 8 --output-dir /tmp/pql_reduced_check
python tests/test_reduced_pql.py
```

## Limitations

The scalar environment is a reduced proxy. It validates PQL mechanism contracts and executable recovery evidence, not the paper's Isaac Gym returns or wall-clock speedups.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…