Skip to content
Back to skills

Ppo Trajectory Advantage Estimation

ASecurity

Compute bootstrapped returns and generalized advantage estimates for fixed-horizon PPO rollout batches.

  • 247 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 9, 2026
toolspythonbash

Works with

  • terminal
  • cli

Security analysis

A100/100

Scanned September 9, 2026

npx -y skills add VectorSpaceLab/AREX-Skill --skill ppo_trajectory_advantage_estimation --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Ppo Trajectory Advantage Estimation?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Ppo Trajectory Advantage Estimation
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vectorspacelab-ppo-trajectory-advantage-estimation/badge)](https://www.skillsdirectory.com/skills/vectorspacelab-ppo-trajectory-advantage-estimation)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

SKILL.md
---
name: ppo_trajectory_advantage_estimation
description: Compute bootstrapped returns and generalized advantage estimates for fixed-horizon PPO rollout batches.
---

# PPO Trajectory Advantage Estimation

## When To Use

Use this skill when a PPO recovery or implementation has a fixed-horizon rollout segment and needs advantages and returns before optimizing the surrogate objective.

Do not use it to update policy parameters or evaluate the clipped surrogate directly.

## Inputs

- `rewards`: non-empty numeric reward sequence.
- `value_predictions`: value estimate for each reward timestep.
- `terminal_flags`: booleans or 0/1 markers where true means the transition ended an episode.
- `next_value_prediction`: bootstrap value after the segment.
- `gamma`: discount factor.
- `gae_lambda`: trace parameter.

## Outputs

JSON containing `advantages`, `returns`, `normalized_advantages`, and summary statistics.

## Workflow

1. Validate rollout arrays have equal length.
2. Append the bootstrap value to the value sequence.
3. Iterate backward over the segment.
4. Mask continuation at terminal boundaries.
5. Compute `delta = reward + gamma * next_value * nonterminal - value`.
6. Compute `gae = delta + gamma * lambda * nonterminal * next_gae`.
7. Return `returns = advantages + values` and safely normalized advantages.

## Validation

Run:

```bash
python tests/test_advantage_estimation.py
python scripts/estimate_advantages.py --rewards '[1,1]' --values '[0.5,0.25]' --terminals '[false,true]' --next-value 0 --gamma 0.99 --gae-lambda 0.95
```

## Limitations

The skill assumes rewards and values have already been sampled by a policy. It does not construct environments or perform policy optimization.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…