Skip to content
Back to skills

Go Explore Robustification Evaluation

ASecurity

Evaluate Go-Explore archived trajectories with deterministic replay and bounded perturbation checks for recovery evidence.

  • 247 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 9, 2026
researchpythongo

Security analysis

A100/100

Scanned September 9, 2026

npx -y skills add VectorSpaceLab/AREX-Skill --skill go_explore_robustification_evaluation --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Go Explore Robustification Evaluation?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Go Explore Robustification Evaluation
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vectorspacelab-go-explore-robustification-evaluation/badge)](https://www.skillsdirectory.com/skills/vectorspacelab-go-explore-robustification-evaluation)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

SKILL.md
---
name: go_explore_robustification_evaluation
description: Evaluate Go-Explore archived trajectories with deterministic replay and bounded perturbation checks for recovery evidence.
---

# Go-Explore Robustification Evaluation

Use this skill after a Go-Explore Phase 1 archive has produced one or more promising trajectories. It separates discovery from evaluation by replaying the discovered trajectory and reporting numeric success metrics. In bounded recovery, it provides a reduced proxy for the paper's robustification phase.

## Inputs

- Best archived trajectory with action list and target score.
- Environment factory or compatible step function.
- Perturbation policy or deterministic replay setting.

## Outputs

- Success rate, score, and replay trace.
- Mechanism checks identifying whether this is full robustification or a reduced proxy.
- Validator-ready metric fields for recovery artifacts.

## Workflow

1. Replay the archived trajectory in a fresh environment from the start state.
2. Confirm that the replay reaches the expected goal or score.
3. Optionally run bounded perturbation checks by starting from nearby curriculum points or injecting no-op perturbations.
4. Report numeric metrics and declare reduced/proxy status when no neural policy is trained.

## Validation

Run `python tests/test_robustification.py` or validate with `validate_skill_tree.py --run-tests`.

## Limitations

This skill does not claim to reproduce the paper's multi-day neural robustification. It creates auditable reduced evidence that discovery and robustness checks are distinct.

Refinement cycle 3 note: bounded perturbation replay must be logged separately from deterministic replay; success rate is treated as proxy evidence, not full robustification.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…