Skip to content
Back to skills

Positive Advantage Sil Loss

ASecurity

Compute Self-Imitation Learning policy and value losses using the paper's positive-advantage gate.

  • 247 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 9, 2026
testingpythonbash

Security analysis

A100/100

Scanned September 9, 2026

npx -y skills add VectorSpaceLab/AREX-Skill --skill positive_advantage_sil_loss --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Positive Advantage Sil Loss?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Positive Advantage Sil Loss
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vectorspacelab-positive-advantage-sil-loss/badge)](https://www.skillsdirectory.com/skills/vectorspacelab-positive-advantage-sil-loss)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

SKILL.md
---
name: positive_advantage_sil_loss
description: Compute Self-Imitation Learning policy and value losses using the paper's positive-advantage gate.
---

# Positive-Advantage SIL Loss

Use this skill when a recovery or implementation needs the exact SIL objective from Equations 1-3 of "Self-Imitation Learning". It is suitable for checking full actor-critic implementations or reduced scalar proxies.

Do not use this skill for ordinary actor-critic advantages that allow negative policy pressure; SIL clamps negative advantages to zero.

## Inputs
- `returns`: Monte Carlo returns from replay records.
- `values`: current value estimates for the same states.
- `action_probabilities` or selected-action probabilities from the policy.
- `beta`: value-loss coefficient.

## Outputs
- Positive advantages.
- Valid-sample mask/count.
- Policy loss, value loss, total loss.

## Workflow
1. Compute `advantage = return - value`.
2. Clamp to `positive_advantage = max(advantage, 0)`.
3. Compute `policy_loss = mean(-log(prob_action) * positive_advantage)`.
4. Compute `value_loss = mean(0.5 * positive_advantage^2)`.
5. Return `policy_loss + beta * value_loss` and diagnostics.

## Validation
Run:

```bash
python tests/test_sil_loss.py
```

The test verifies exact loss values and zero contribution for non-positive advantages.

## Limitations
- This script computes scalar diagnostics, not automatic differentiation.
- Callers are responsible for mapping full policy distributions to selected-action probabilities.

## Zero-Pressure Ablation
If every replay return is less than or equal to the current value estimate, SIL must report zero valid samples and apply no imitation pressure. Treat this as expected behavior, not a metric failure.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…