Skip to content
Back to skills

Ppo Actor Critic Update Loop

ASecurity

Execute a reduced PPO actor-critic update with clipped surrogate, value loss, and optimizer-step evidence.

  • 247 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 9, 2026
toolspython

Works with

  • cli

Security analysis

A100/100

Scanned September 9, 2026

npx -y skills add VectorSpaceLab/AREX-Skill --skill ppo_actor_critic_update_loop --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Ppo Actor Critic Update Loop?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Ppo Actor Critic Update Loop
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vectorspacelab-ppo-actor-critic-update-loop/badge)](https://www.skillsdirectory.com/skills/vectorspacelab-ppo-actor-critic-update-loop)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

SKILL.md
---
name: ppo_actor_critic_update_loop
description: Execute a reduced PPO actor-critic update with clipped surrogate, value loss, and optimizer-step evidence.
---

# PPO Actor-Critic Update Loop

Use this skill to build bounded PPO recovery experiments or audits that need executable evidence of a policy/value update. It is intended for mechanism-faithful reduced recovery and small deterministic tests, not for claiming full MuJoCo or Atari reproduction.

## Inputs

- Rollout steps and `last_value`, or a precomputed batch with advantages and returns.
- Old log probabilities and action features for a tiny trainable policy-ratio model.
- Scalar value predictions and learning-rate/loss coefficients.
- Paths to the rollout and clipped-objective skill scripts when composing generated skills.

## Outputs

- `loss_before` and `loss_after`.
- `params_before` and `params_after` for validator-compatible optimizer evidence.
- Mechanism diagnostics for GAE execution, clipped objective execution, and optimizer execution.

## Workflow

1. Compute advantages and returns from rollout steps.
2. Compute the clipped surrogate objective from old and current log probabilities.
3. Add a value-function squared-error term and optional entropy proxy.
4. Estimate deterministic finite-difference gradients for a tiny scalar parameterization.
5. Apply at least one optimizer step.
6. Recompute loss and write a training trace.

## Validation

Run `python tests/test_update.py` from this skill directory. The tests assert parameter changes, finite losses, and execution of GAE and clipped-surrogate paths.

## Limitations

The included optimizer is intentionally tiny and deterministic. It validates PPO's update mechanics under soft-mode recovery but is not a replacement for large-scale neural-network training.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…