Skip to content
Back to skills

Wmdp Eval

ASecurity

Audit a dual-use capability claim against WMDP, unlearning evaluation, and elicitation studies. Use when you need help with wmdp eval.

  • 8 stars
  • 0 votes
  • 0 copies
  • 3 views
  • Added September 8, 2026
researchsecurity

Security analysis

A100/100

Scanned September 8, 2026

npx -y skills add anubhavg-icpl/vibe --skill wmdp-eval --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Wmdp Eval?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Wmdp Eval
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/anubhavg-icpl-wmdp-eval/badge)](https://www.skillsdirectory.com/skills/anubhavg-icpl-wmdp-eval)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: wmdp-eval
description: Audit a dual-use capability claim against WMDP, unlearning evaluation, and elicitation studies. Use when you need help with wmdp eval.
license: CC-BY-NC-SA-4.0
phase: 18
lesson: 17
metadata:
  version: 1.0.0
  tags: [wmdp, rmu, dual-use, biosecurity, cybersecurity, chemistry]
---

Given a dual-use capability claim ("our model does not meaningfully help with bioweapons / cyberattack / chemistry"), audit the supporting evaluation.

Produce:

1. Benchmark coverage. Was WMDP (or an equivalent yellow-zone benchmark) run? Report per-domain scores (bio, cyber, chem). A claim without per-domain numbers cannot be evaluated.
2. Unlearning trace. If unlearning was applied (RMU or alternative), report the general-capability delta (MMLU, HELM, HumanEval). Unlearning with no general-capability report is not credible.
3. Refusal-path-audit. Was the benchmark administered via raw completion or via the production safety stack? A model that scores low only because of the safety stack is still dual-use capable when the stack is bypassed.
4. Elicitation study. Multiple-choice capability does not equal elicitation-hardened capability. Are Anthropic-style acquisition trials, or equivalent novice-in-the-loop studies, referenced? If not, the claim is limited to WMDP-style evidence.
5. Novice-vs-expert split. Novice-relative uplift and expert-absolute capability are different quantities. Are both addressed?

Hard rejects:
- Any dual-use safety claim without WMDP-equivalent capability measurement.
- Any unlearning claim without general-capability delta.
- Any "no meaningful uplift" claim without novice-in-the-loop study.

Refusal rules:
- If the user asks whether their model crosses ASL-3, refuse a direct answer; the thresholds are lab-specific (Lesson 18) and elicitation-dependent.
- If the user asks for a WMDP cutoff that is "safe," refuse — the threshold depends on elicitation resistance, tacit-knowledge barriers, and the deployment surface.

Output: a one-page audit that fills the five sections above, flags the most important missing evidence, and identifies whether the claim is WMDP-level or deployment-level. Cite Li et al. (arXiv:2403.03218) once as the benchmark source.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…