Skip to content
Back to skills

Policy

ASecurity

Drive PWN::AI::Agent::Policy from pwn_eval.

  • 86 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 10, 2026
securityrustrubyshellgit

Works with

  • terminal

Security analysis

A100/100

Scanned September 10, 2026

npx -y skills add 0dayInc/pwn --skill policy --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Policy?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Policy
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/0dayinc-policy/badge)](https://www.skillsdirectory.com/skills/0dayinc-policy)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: pwn-ai-agent-policy
description: Drive PWN::AI::Agent::Policy from pwn_eval.
license: MIT
allowed-tools: [pwn, pwn_eval]
metadata:
  bundled: true
  generated: true
  module: PWN::AI::Agent::Policy
  source: pwn/ai/agent/policy.rb
---

# PWN::AI::Agent::Policy

PWN::AI::Agent::Policy is the LIVE tabular RL controller that pwn-ai did not have before R5. Everything else in the harness is retrieval-plus-policy: scores are written to disk and re-injected as prose, or exported later for optional LoRA. This module is the missing MDP: state s — discretized (kind, task, plan, completeness, usable, last, fail) action a — tool name, or "final" reward r — step: 0 (spam cost −0.01 after 8 tools); terminal: judge × confidence next s' — state after the tool result Each Loop turn is one episode. Trusted environment/prerequisite bins scope fallback history; independently evidenced terminal attribution uses isolated action targets, not future-return credit for busywork. Transitions land in ~/.pwn/policy_traj.jsonl. Q(s,a) and REINFORCE logits H(s,a) are updated from those tuples and persisted in ~/.pwn/policy.json. The learned Q values are an ADVISORY term in Registry.rank. They never replace TaskSummarizer planning, plan_first, or CORE_TOOLS. Disable with PWN::Env[:ai][:agent][:policy] = false.

## When to use

Call `PWN::AI::Agent::Policy` from `pwn_eval` when the task needs this module.
Do not reimplement it in shell.

## Methodologies

Generated from `pwn/ai/agent/policy.rb`. Prefer the public class methods below.
Class methods take `(opts = {})` and read `opts`.

## How to call

```ruby
PWN::AI::Agent::Policy.help
PWN::AI::Agent::Policy.state(opts)
```

## Public methods

- `state`
- `observed_state`
- `observed_context`
- `cold`
- `warm`
- `episode_budget_met`
- `begin_episode`
- `observe_step`
- `finish`
- `update_q`
- `update_pg`
- `q`
- `value`
- `advantage`
- `recommend`
- `current_state`
- `current_episode`
- `current_context_state`
- `detach_episode`
- `attach_episode`
- `load`
- `save`
- `trajectories`
- `stats`
- `evaluate`
- `to_context`
- `lean`
- `reset`
- `enabled`
- `authors`
- `help`
- `warmup`
- `maybe_warmup`
- `attach_episode!`
- `cold?`
- `detach_episode!`
- `enabled?`
- `episode_budget_met?`
- `lean!`
- `maybe_warmup!`
- `update_pg!`
- `update_q!`
- `warm?`
- `warmup!`

## Source

`pwn/ai/agent/policy.rb`

## Verification

`PWN::AI::Agent::Policy.respond_to?(:state)` after the
module is loaded. Read the source for parameter names.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…