Skip to content
Back to skills

Ppo Clipped Surrogate Objective

ASecurity

Compute and validate PPO clipped probability-ratio surrogate objectives and trust-region-style diagnostics.

  • 247 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 9, 2026
developmentpythonrust

Works with

  • cli

Security analysis

A100/100

Scanned September 9, 2026

npx -y skills add VectorSpaceLab/AREX-Skill --skill ppo_clipped_surrogate_objective --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Ppo Clipped Surrogate Objective?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Ppo Clipped Surrogate Objective
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vectorspacelab-ppo-clipped-surrogate-objective/badge)](https://www.skillsdirectory.com/skills/vectorspacelab-ppo-clipped-surrogate-objective)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

SKILL.md
---
name: ppo_clipped_surrogate_objective
description: Compute and validate PPO clipped probability-ratio surrogate objectives and trust-region-style diagnostics.
---

# PPO Clipped Surrogate Objective

Use this skill when a PPO implementation, recovery harness, or audit needs the paper's clipped surrogate objective from Equation 7. Do not use it as a generic ratio clamp: the selected term must be the minimum of the unclipped and clipped advantage-weighted objectives.

## Inputs

- `new_log_probs`: log probabilities under the updated/current policy.
- `old_log_probs`: log probabilities under the behavior policy.
- `advantages`: advantage estimates aligned with the sampled actions.
- `clip_epsilon`: clipping half-width, often `0.2`.

## Outputs

- Per-sample ratios, unclipped terms, clipped terms, and selected terms.
- Mean objective and minimization loss.
- Clip fraction and approximate KL diagnostics.

## Workflow

1. Compute `ratio = exp(new_log_prob - old_log_prob)`.
2. Compute `ratio * advantage` for the conservative-policy-iteration surrogate term.
3. Clamp the ratio to `[1 - epsilon, 1 + epsilon]` and multiply by the same advantage.
4. Select the smaller objective contribution per sample.
5. Average selected terms and negate for minimization.
6. Record diagnostics for recovery analysis.

## Validation

Run `python tests/test_surrogate.py` from this skill directory. The tests cover positive-advantage upper clipping, negative-advantage lower clipping, and in-range no-op behavior.

## Limitations

This skill computes the policy objective only. Value loss, entropy bonus, rollout generation, and optimizer execution belong to separate modules.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…