Skip to content
Back to skills

Score Function Estimator

ASecurity

Compute REINFORCE score-function update terms for sampled stochastic policy actions with scalar rewards and baselines.

  • 247 stars
  • 0 votes
  • 0 copies
  • 3 views
  • Added September 9, 2026
testingpythongobashgit

Security analysis

A100/100

Scanned September 9, 2026

npx -y skills add VectorSpaceLab/AREX-Skill --skill score_function_estimator --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Score Function Estimator?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Score Function Estimator
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vectorspacelab-score-function-estimator/badge)](https://www.skillsdirectory.com/skills/vectorspacelab-score-function-estimator)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

SKILL.md
---
name: score_function_estimator
description: Compute REINFORCE score-function update terms for sampled stochastic policy actions with scalar rewards and baselines.
---

# Score Function Estimator

Use this skill when a recovery or implementation needs the local REINFORCE likelihood-ratio primitive for a sampled stochastic action. Do not use it for deterministic argmax policies, supervised cross-entropy labels, or value-function-only updates.

## Inputs

- Bernoulli action probability `prob_action_one` in `(0, 1)`.
- Sampled action `0` or `1`.
- Scalar reward or return.
- Optional action-independent baseline.

## Outputs

- `grad_log_prob`: derivative of the sampled action log probability with respect to the Bernoulli logit.
- `advantage`: `reward - baseline`.
- `update`: `advantage * grad_log_prob`.

## Workflow

1. Confirm the action was sampled from the same stochastic policy whose probability is supplied.
2. Clamp the probability only for numerical safety; do not change the action.
3. Compute `grad_log_prob = action - prob_action_one` for a Bernoulli logit policy.
4. Center reinforcement with an action-independent baseline if provided.
5. Return the update direction, leaving learning-rate scaling to the training loop.

## Validation

Run:

```bash
python scripts/score_estimator.py --prob 0.4 --action 1 --reward 1.0 --baseline 0.2
python -m pytest tests
```

The tests verify update signs and that a baseline changes magnitude but preserves the ascent direction in a high-reward action example.

## Limitations

This skill implements the Bernoulli-logit case used by the bounded recovery. Categorical policies use the same score-function principle but require vector log-probability gradients.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…