Skip to content
Back to skills

W2sg Pgr

ASecurity

Audit a scalable-oversight or W2SG claim via the performance-gap-recovered metric. Use when you need help with w2sg pgr.

  • 8 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 8, 2026
developmentrustgoperformance

Security analysis

A100/100

Scanned September 8, 2026

npx -y skills add anubhavg-icpl/vibe --skill w2sg-pgr --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of W2sg Pgr?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for W2sg Pgr
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/anubhavg-icpl-w2sg-pgr/badge)](https://www.skillsdirectory.com/skills/anubhavg-icpl-w2sg-pgr)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: w2sg-pgr
description: Audit a scalable-oversight or W2SG claim via the performance-gap-recovered metric. Use when you need help with w2sg pgr.
license: CC-BY-NC-SA-4.0
phase: 18
lesson: 11
metadata:
  version: 1.0.0
  tags: [scalable-oversight, weak-to-strong, pgr, debate, recursive-reward-modeling]
---

Given a scalable-oversight or W2SG paper / report, audit whether the setup supports its claim.

Produce:

1. Weak / strong identification. Explicitly name the weak supervisor and the strong model. Is the capability gap measured in parameters, training tokens, benchmark score, or task-specific evaluation?
2. Ceiling definition. What is the strong model's supervised ceiling on the task? Without a ceiling, PGR cannot be computed.
3. PGR computation. PGR = (fine-tuned - weak) / (ceiling - weak). Check sign, magnitude, and denominator. Small denominators inflate PGR artificially.
4. Prior-leakage check. Does the strong model's pre-training data include the task's ground truth? If yes, "recovery" may be prior retrieval rather than generalization.
5. Alignment-vs-capability split. Is the weak-to-strong gap a capability gap or an alignment gap? Burns et al. 2023 is explicit that their gap is capability-shaped; alignment-shaped gaps may behave differently.

For scalable-oversight mechanism audits:
- Debate: identify the judge's knowledge, the debater structure, and whether the task rewards truth-leans. Cite Khan et al. 2024 (arXiv:2402.06782) on where debate helps and fails.
- RRM: identify the recursion depth and what happens if U+1 is already untrustworthy.
- Task decomposition: identify the decomposition procedure and whether sub-tasks are independently checkable.

Hard rejects:
- Any PGR claim without a ceiling on gold labels.
- Any W2SG claim that claims to solve alignment — W2SG measures capability recovery, not alignment.
- Any debate-mechanism claim that ignores the 2024 empirical literature on when debate helps vs hurts.

Refusal rules:
- If the user asks "does W2SG solve superalignment," refuse the binary answer and explain PGR is a measurable, not a solution.
- If the user asks which scalable-oversight mechanism is best, refuse — the answer is task-dependent.

Output: a one-page audit that fills the five sections above, reports or requests PGR, and flags whether the weak-strong gap is capability-shaped or alignment-shaped. Cite Burns et al. 2023 and Lang et al. (arXiv:2501.13124) once each.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…