Skip to content
Back to skills

Agent Tool Safety Guard

ASecurity

Design or review least-privilege tool and function access for a large language model (LLM) agent, containing excessive agency and tool misuse. Covers Open Worldwide Application Security Project (OWASP) identifiers LLM03 (excessive agency), ASI02 (tool misuse), and the tool-enabled ASI05 code-execution slice. Build a per-tool permission matrix, validate arguments before execution, use the calling user's authority, gate high-impact actions behind human approval, and map tool-chain abuse. Code-e...

  • 4 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 11, 2026
ai-agentspythonrustgoshellgitapibackendsecurity

Works with

  • cli
  • api

Security analysis

A100/100

Pro scans all 4 files and shows the line behind each finding

Scanned October 5, 2026

npx -y skills add ModernNomad-98/Project-Aegis --skill agent-tool-safety-guard --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Agent Tool Safety Guard?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Agent Tool Safety Guard
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/modernnomad-98-agent-tool-safety-guard/badge)](https://www.skillsdirectory.com/skills/modernnomad-98-agent-tool-safety-guard)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: agent-tool-safety-guard
description: Design or review least-privilege tool and function access for a large language model (LLM) agent, containing excessive agency and tool misuse. Covers Open Worldwide Application Security Project (OWASP) identifiers LLM03 (excessive agency), ASI02 (tool misuse), and the tool-enabled ASI05 code-execution slice. Build a per-tool permission matrix, validate arguments before execution, use the calling user's authority, gate high-impact actions behind human approval, and map tool-chain abuse. Code-execution tools need a sandbox and approval. Composes human-approval-boundary and agent-authorization-matrix; ai-human-in-the-loop-designer designs the product review workflow behind an approval gate. Use when an agent can call tools, functions or application programming interfaces (APIs). Do NOT use for injection defense, sandbox design, retrieval authorization, what tool schemas reveal to users (hidden-context-exposure-reviewer), or standing merge/deploy authority.
---

# Agent Tool Safety Guard

**Reading key:** A large language model (LLM) agent can request tool calls.
Open Worldwide Application Security Project (OWASP) code LLM03 (2026 edition;
LLM06 in 2025) names excessive agency; codes ASI02 and ASI05 from the OWASP
Top 10 for Agentic Applications (2026) refer here to tool misuse and
tool-enabled code execution. An application programming interface (API) is a
software call boundary; natural language (NL) is ordinary user text.
`authz` means authorization, and remote code execution (RCE) means an
attacker can cause code to run across a trust boundary.

## Purpose

Contain excessive agency (LLM03): design or review the tool/function surface
of an LLM agent so it can only do what the task needs, with the authority of
the user it acts for, and cannot cause irreversible harm without a human. The
deliverable is a per-tool permission matrix (side effects, blast radius,
identity), argument validation before execution, an identity/authority model
that binds tool calls to the CALLING USER, approval gates on high-impact
actions, and a map of tool-chain composition abuse. Approval mechanics come
from `human-approval-boundary`; standing agent authority from
`agent-authorization-matrix` — this skill composes both.

## Use When

- Use when: an LLM agent, assistant, or workflow can invoke tools, functions,
  plugins, or APIs and you need to scope what it may do and how far a mistake
  or injection can reach.
- Use when: reviewing an agent for excessive agency — too many tools, too
  broad scopes, actions that run as a privileged service account, missing
  approval on destructive operations.
- Use when: adding a new tool to an existing agent and its blast radius needs
  assessing.
- Use when: reviewing tool misuse/exploitation paths (ASI02) — a manipulated
  agent abusing its legitimate tools within granted scopes, side-effect
  limits per tool, or a code-execution tool (interpreter/shell/eval) whose
  blast-radius class and approval posture need setting (the tool-side slice
  of ASI05; the sandbox internals are `llm-output-safety-reviewer`'s).
- Do NOT use when: the concern is stopping injected instructions from
  reaching the tools (`prompt-injection-defender` — this skill assumes the
  boundary and designs it).
- Do NOT use when: the agent EXECUTES generated code
  (`llm-output-safety-reviewer` for the exec/sandbox surface), retrieves
  documents (`rag-security-architect`), or the question is agent-vs-human
  merge/deploy authority (`agent-authorization-matrix`).
- Do NOT use when: the question is what tool names, descriptions, and
  parameter schemas reveal if a user extracts them, or whether a rule written
  in a tool description is being relied on as a gate
  (`hidden-context-exposure-reviewer`, LLM08 Hidden Context Exposure); that
  skill routes the enforcement fix back here.
- Do NOT use when: the ask is the review workflow behind a tool gate, or AI
  output that is not a tool call — that is `ai-human-in-the-loop-designer`.

## Inputs to Inspect

1. The tool/function inventory: every tool the agent can call, its
   description as the model sees it, its parameters, and what it actually does.
2. Side effects per tool: read vs write, reversible vs irreversible, internal
   vs external, money-spending, data-exposing.
3. The identity the tool runs as: the calling user's permissions, the agent's
   own identity, or a shared service account (the last is usually the bug).
4. Argument handling: is input validated/typed before execution, or is the
   model's free-text passed through?
5. Existing approval gates and their triggers; `human-approval-boundary` and
   `agent-authorization-matrix` output where present.
6. The autonomy context: is the agent autonomous, human-in-the-loop, or
   advisory; how tool outputs feed back into the next model call.
7. Code-execution tools if any (interpreter, shell, eval, code-runner):
   their sandbox posture, what executed code can reach, and which
   natural-language paths can trigger execution (ASI05 tool slice — the
   sandbox rubric itself lives with `llm-output-safety-reviewer`).

## Workflow

1. **Inventory tools and side effects.** For each tool, record: purpose, real
   side effect, reversibility, blast radius (one record / one tenant / all
   tenants / external), and cost. Class any code-execution tool (interpreter,
   shell, eval) as maximal blast radius from the start (ASI05). No tool
   surface to inspect → Stop Conditions.
2. **Apply least privilege.** Challenge every tool: does the task require it?
   Can a read-only or narrower-scope version do? Remove or narrow tools that
   exceed the task. Fewer tools and tighter scopes shrink the surface an
   injection or hallucination can reach.
3. **Bind tool calls to the calling user's authority.** Every tool executes
   with the CALLING USER's permissions and tenant scope, enforced by code —
   not the agent's identity and not a broad service account. An agent must not
   be able to do for a user what the user cannot do themselves.
4. **Validate arguments before execution.** Each tool's arguments are checked
   against a schema and value constraints (allowlists, ranges, tenant-scoped
   ids) BEFORE the side effect runs — the model's output is untrusted input
   to the tool (compose `structured-output-validator`).
5. **Gate high-impact actions.** Irreversible, destructive, costly, or
   cross-boundary actions require human approval; define the trigger per tool.
   A product agent's approvals run through the review workflow designed by
   `ai-human-in-the-loop-designer`; a coding agent's own risky steps stop at
   `human-approval-boundary`. Standing "agent may never do X autonomously"
   rules come from `agent-authorization-matrix`. Code-execution tools default
   to sandbox-required plus approval (per-tool side-effect limits set here;
   the sandbox itself per `llm-output-safety-reviewer`, ASI05).
6. **Map tool-chain composition abuse** using
   [references/tool-permission-matrix.md](references/tool-permission-matrix.md):
   where one tool's (untrusted) output becomes another tool's arguments, a
   safe-looking chain can compose into harm (read → summarize → send). Model
   the chains, not just single calls — including natural-language-driven
   execution paths where untrusted content steers WHICH tool runs with WHICH
   arguments (ASI02 misuse runs through legitimate grants).
7. **Design containment and telemetry.** Rate/quantity limits per tool,
   a kill switch to disable a tool or the agent, and logging of every tool
   call with arguments and outcome (compose `observability-operator`).
   Confirmed live abuse routes to the human incident owner, who follows the
   approved response runbook. The `incident-response-runbook` skill can
   author or improve that procedure later; it does not run the incident.
8. **Design the red-team cases.** For each high-risk tool: an injection or
   hallucination that tries to trigger it out of scope, with the expected
   SAFE outcome (denied at authz/approval). Hand to `ai-evaluation-harness`.

## Output Format

```
AGENT TOOL SAFETY — <agent/feature>
Tool matrix:
  <tool> | side effect: <r/w, reversible?> | blast radius: <scope> | runs as: <identity> | cost: <>
Least-privilege changes: <tools removed/narrowed + why>
Identity model: <calling-user authority enforcement per tool>
Argument validation: <schema/constraints per tool> (→ structured-output-validator)
Approval gates: <tool → trigger → approver> (→ human-approval-boundary)
Tool-chain abuse: <chain → composed harm → break point>
Containment: <rate/qty limits | kill switch | telemetry> (→ observability-operator)
Red-team cases: <out-of-scope trigger attempt → SAFE outcome> (→ ai-evaluation-harness)
Residual risk: <what remains + named acceptor>
```

## Validation Checklist

- [ ] Every tool has a recorded side effect, reversibility, blast radius, and
      the identity it runs as.
- [ ] Least privilege applied: unnecessary or over-broad tools removed or
      narrowed with rationale.
- [ ] Every tool executes with the calling user's authority/tenant scope,
      enforced by code — no broad service-account execution left unjustified.
- [ ] Arguments are schema/constraint-validated before the side effect runs.
- [ ] Irreversible/costly/cross-boundary actions are approval-gated with a
      defined trigger.
- [ ] Tool-chain composition abuse paths are mapped, not just single calls.
- [ ] Code-execution tools (if any) are classed maximal-blast-radius:
      sandbox-required (per `llm-output-safety-reviewer`), approval-gated,
      with misuse and NL-driven execution paths mapped (ASI02/ASI05).
- [ ] Kill switch, per-tool limits, and per-call telemetry are specified;
      live abuse routes to the human incident owner and approved runbook.

## Tool Permission Rules

- Least privilege by default: a tool the task doesn't need is removed, not
  left "just in case".
- Tools run as the user, not the agent: an agent must never let a user do
  through it what they can't do directly (confused-deputy prevention).
- Model output driving a tool is untrusted input: validate arguments before
  execution; never pass free-text straight to a side effect.
- Irreversible actions are approval-gated — the model does not get to decide
  to delete, pay, deploy, or externally send on its own.
- The tool description the model sees is part of the attack surface: a tool
  named/described to invite misuse is a finding.

## Gotchas

- The service-account trap: wiring every tool to one privileged backend
  identity means a single injection can act as admin for all tenants. Bind to
  the caller.
- Over-tooling: giving the agent 30 tools "for flexibility" hands an attacker
  30 primitives. Scope to the task.
- Reversibility is a spectrum: "send email" and "delete account" are both
  writes but only one is catastrophic — gate by impact, not just by verb.
- Chains hide harm: read-file + post-webhook are each innocuous; together
  they exfiltrate. Enumerate compositions.
- Approval fatigue undermines gates: if everything prompts, users click
  through (see the agentic `human-agent-trust-reviewer`). Gate the genuinely
  high-impact actions, not everything.
- "The model decided" is not authorization — authorization is a code
  decision about the user, checked before the effect.
- A code-execution tool makes the rest of the matrix moot if unsandboxed:
  "run this Python" with network and ambient credentials is RCE with extra
  steps — class interpreter/shell/eval tools as their own maximal-blast-
  radius row (ASI05), sandbox-required and approval-gated, never wired to a
  service account. ASI02 misuse needs no new tools at all: the abuse runs
  THROUGH legitimate grants, which is why side-effect limits and chain
  mapping matter even when every single tool looks properly scoped.

## Stop Conditions

- No tool/function inventory or agent design is available — stop; this skill
  scopes concrete tools, not a hypothetical agent.
- A tool runs as a broad service account with no way to bind it to the caller
  and it performs cross-tenant writes — flag as a blocking finding and route
  the fix through `human-approval-boundary`.
- The gap is really injection reaching the tools, code execution, or standing
  merge/deploy authority — hand to the owning skill.
- A review finds a tool already being abused in production — route to the
  human incident owner and approved response runbook. Activating a kill
  switch or revoking access is a separately authorized live action; use
  `incident-response-runbook` only to author or improve the procedure.

## Supporting Files

- [references/tool-permission-matrix.md](references/tool-permission-matrix.md)
  — the per-tool matrix template, blast-radius rubric, calling-user identity
  patterns, argument-validation checklist, the tool-chain composition
  abuse catalog, and the ASI02/ASI05 extension (tool misuse paths,
  code-execution tool class, NL-driven execution).
- `evals/evals.json` — trigger + behavior cases.
- `evals/trigger-evals.json` — discrimination within the output & agency
  cluster and against `agent-authorization-matrix` and `human-approval-boundary`.

Files in this skill

  • SKILL.md11.5 KB
  • evals/evals.json3.2 KB
  • evals/trigger-evals.json3.1 KB
  • references/tool-permission-matrix.md6.8 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…