Skip to content
Back to skills

Sample Factory Pbt Self Play

ASecurity

Specify Sample Factory-style multi-policy self-play assignments and population-based exploit/mutate decisions.

  • 247 stars
  • 0 votes
  • 0 copies
  • 3 views
  • Added September 9, 2026
businesspythonbash

Works with

  • cli

Security analysis

A100/100

Scanned September 9, 2026

npx -y skills add VectorSpaceLab/AREX-Skill --skill sample_factory_pbt_self_play --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Sample Factory Pbt Self Play?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Sample Factory Pbt Self Play
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vectorspacelab-sample-factory-pbt-self-play/badge)](https://www.skillsdirectory.com/skills/vectorspacelab-sample-factory-pbt-self-play)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

SKILL.md
---
name: sample_factory_pbt_self_play
description: Specify Sample Factory-style multi-policy self-play assignments and population-based exploit/mutate decisions.
---

# Sample Factory PBT Self-Play

Use this skill when planning or auditing Sample Factory-style self-play experiments with multiple policies and population-based training.

Do not use it for single-policy throughput recovery unless the recovery explicitly validates self-play or population management.

## Inputs
- Policy ids and hyperparameters.
- Episode agent ids and available policy ids.
- Policy scores or win rates.
- Replacement threshold and mutation factors.

## Outputs
- Episode policy-assignment log.
- Ranked population summary.
- Exploit/mutate decisions.
- Mutated hyperparameters clipped to bounds.

## Workflow
1. Assign each agent in an episode to a policy id while keeping rollout workers policy-agnostic.
2. Rank policies by win rate or task score at the mutation interval.
3. Mark weak policies below a threshold fraction of the best policy for replacement.
4. Copy hyperparameters from the best policy to replaced policies.
5. Apply deterministic or seeded multiplicative mutations within configured bounds.
6. Record decisions separately from environment rewards and model updates.

## Validation
Run:

```bash
python scripts/pbt_self_play.py
python tests/test_pbt_self_play.py
```

## Limitations
This skill models population-control decisions. It does not train policies or simulate VizDoom matches.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…