Skip to content
Back to skills

Jsrl Value Exploration Update

ASecurity

Apply a minimal value-based exploration-policy update to Jump-Start Reinforcement Learning trajectories and record optimizer evidence.

  • 247 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 9, 2026
toolspythonbash

Works with

  • terminal

Security analysis

A100/100

Scanned September 9, 2026

npx -y skills add VectorSpaceLab/AREX-Skill --skill jsrl_value_exploration_update --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Jsrl Value Exploration Update?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Jsrl Value Exploration Update
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vectorspacelab-jsrl-value-exploration-update/badge)](https://www.skillsdirectory.com/skills/vectorspacelab-jsrl-value-exploration-update)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

SKILL.md
---
name: jsrl_value_exploration_update
description: Apply a minimal value-based exploration-policy update to Jump-Start Reinforcement Learning trajectories and record optimizer evidence.
---

# JSRL Value Exploration Update

## When To Use

Use this skill after JSRL switching rollouts have produced trajectories and the recovery needs to show a real value-based policy update. It is suitable for reduced proxy runs and for sanity-checking full implementations.

## Inputs

- Transitions containing `state`, `action`, `reward`, `next_state`, and `done`.
- Q-value table or parameter dictionary for the exploration-policy.
- Learning rate and discount factor.

## Outputs

- Updated Q-value parameters.
- Training trace with `loss_before`, `loss_after`, `params_before`, `params_after`, and `optimizer_state_changed`.
- Greedy action helper for the exploration-policy.

## Workflow

1. Build a deterministic key for each state-action value.
2. Compute TD targets with terminal-state handling.
3. Record squared TD loss before the update.
4. Apply Q-learning style updates to trainable parameters.
5. Recompute loss and save parameter-change evidence.
6. Use the updated greedy policy in later exploration rollouts.

## Validation

Run:

```bash
python scripts/value_update.py --demo
python -m pytest tests
```

The demo prints before/after Q parameters and a validator-compatible training trace.

## Limitations

This reduced update is not a replacement for IQL, DQN, or QT-Opt. It preserves the paper-relevant requirement that JSRL data feeds a value-based train-policy step with auditable parameter changes.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…