Skip to content
Back to skills

Qdagger Loss

ASecurity

Compute the QDagger objective for policy-to-value reincarnating RL using n-step TD loss plus teacher-policy distillation.

  • 247 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 9, 2026
developmentpythonbash

Security analysis

A100/100

Scanned September 9, 2026

npx -y skills add VectorSpaceLab/AREX-Skill --skill qdagger_loss --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Qdagger Loss?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Qdagger Loss
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vectorspacelab-qdagger-loss/badge)](https://www.skillsdirectory.com/skills/vectorspacelab-qdagger-loss)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

SKILL.md
---
name: qdagger_loss
description: Compute the QDagger objective for policy-to-value reincarnating RL using n-step TD loss plus teacher-policy distillation.
---

# QDagger Loss

## When To Use

Use this skill when implementing or validating a reduced or full QDagger-style policy-to-value reincarnating RL experiment. It is appropriate when you have transition records, student Q-values, and teacher action probabilities. Do not use it as a standalone imitation-only objective; the TD and distillation terms must remain separately visible.

## Inputs

- `q_values`: mapping from state ids to action-value lists.
- `transitions`: records with `state`, `action`, `n_step_return`, `discount`, `next_max_q`, and `teacher_policy`.
- `temperature`: positive softmax temperature for deriving the student policy from Q-values.
- `lambda_t`: current distillation coefficient.

## Outputs

- `td_loss`: mean squared Bellman error.
- `distillation_loss`: mean teacher cross-entropy against the student softmax policy.
- `combined_loss`: `td_loss + lambda_t * distillation_loss`.
- `examples`: per-transition targets, predictions, and loss terms.

## Workflow

1. Validate that every transition has a known state, valid action index, and normalized teacher policy.
2. Compute each n-step TD target as `n_step_return + discount * next_max_q`.
3. Compute mean squared TD loss over selected actions.
4. Convert each student Q-vector into `softmax(q / temperature)`.
5. Compute teacher-policy cross-entropy.
6. Return both component losses and the combined QDagger loss.

## Validation

Run:

```bash
python scripts/qdagger_loss.py --self-test
python tests/test_qdagger_loss.py
```

## Limitations

This implementation is deterministic and dependency-free. It is intended for recovery harnesses, tests, and small tabular/scalar checks, not high-throughput deep RL training.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…