Skip to content
Back to skills

Common Llm Security

ASecurity

Assumption: the agent can read arbitrary files and delete files based on model decisions, with no human confirmation or equivalent host-enforced policy. **Overall: 🔴 P0 — LLM06 Excessive Agency.** Autonomous deletion is a confirmed high-impact capability. Security score is capped at **40/100** until fixed. | OWASP ID | Status | Review | |---|---|---| | LLM01 Prompt Injection | ⚠️ Needs review | Check whether file contents or user input are concatenated into system prompts. Keep untrusted con...

  • 571 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 24, 2026
developmentrustapisecurity

Works with

  • api

Security analysis

A100/100

Pro scans all 20 files and shows the line behind each finding

Scanned September 24, 2026

npx -y skills add HoangNguyen0403/agent-skills-standard --skill common-llm-security --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Common Llm Security?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Common Llm Security
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/hoangnguyen0403-common-llm-security-82baaa1d/badge)](https://www.skillsdirectory.com/skills/hoangnguyen0403-common-llm-security-82baaa1d)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
Assumption: the agent can read arbitrary files and delete files based on model decisions, with no human confirmation or equivalent host-enforced policy.

**Overall: 🔴 P0 — LLM06 Excessive Agency.** Autonomous deletion is a confirmed high-impact capability. Security score is capped at **40/100** until fixed.

| OWASP ID | Status | Review |
|---|---|---|
| LLM01 Prompt Injection | ⚠️ Needs review | Check whether file contents or user input are concatenated into system prompts. Keep untrusted content in a separate user/data message; do not treat file text as instructions. |
| LLM02 Sensitive Information Disclosure | ⚠️ Needs review | File reads may expose credentials or PII through prompts, responses, or logs. Minimize access and redact sensitive data before logging or persistence. |
| LLM03 Supply Chain | ⚠️ Needs review | Review model, plugins, and packages. Pin source revisions and hashes; hashes provide integrity, not trusted authorship. |
| LLM04 Data & Model Poisoning | ⚠️ Needs review | Treat files, logs, retrieved documents, and proposed learning entries as untrusted. Validate and sanitize data before persistence or indexing. |
| LLM05 Improper Output Handling | ⚠️ Needs review | Never use raw model output directly as a filesystem path or delete command. Parse against a strict schema, canonicalize paths, validate authorization, and reject traversal or ambiguous paths. |
| LLM06 Excessive Agency | 🔴 Confirmed | Read/delete tools operate without confirmation. Require human-in-the-loop approval for every deletion, or restrict deletion to an explicitly authorized, host-enforced sandbox with deny-by-default permissions. |
| LLM07 System Prompt Leakage | ⚠️ Needs review | Ensure tool errors, file contents, and API responses cannot reveal system prompts, policies, credentials, or hidden tool instructions. |
| LLM08 Vector & Embedding Weaknesses | ⚠️ Needs review | If file contents are embedded, sanitize untrusted text and enforce tenant/user namespace isolation. |
| LLM09 Misinformation | ⚠️ Needs review | Do not let an unverified model judgment trigger irreversible deletion, especially for legal, financial, medical, or compliance records. |
| LLM10 Unbounded Consumption | ⚠️ Needs review | Enforce `max_tokens`, invocation rate limits, file-size/read limits, timeout limits, and a maximum agent iteration/depth cap. |

Minimum remediation:

1. Default all file operations to read-only.
2. Require explicit confirmation containing the exact resolved path, reason, and irreversible effect before deletion.
3. Enforce permissions in the host/tool layer; skill text or `allowed-tools` declarations are not sufficient.
4. Use an allowlist of directories, deny symlinks and traversal, canonicalize paths, and restrict deletion to a sandbox.
5. Add dry-run mode, immutable audit logs, rollback/recycle-bin behavior where possible, and emergency disablement.
6. Validate structured tool arguments and sanitize model-generated text before filesystem use.
7. Add tests for prompt injection through file contents, path traversal, symlink escapes, sensitive-file access, repeated deletion loops, and confirmation bypasses.

Files in this skill

  • eval-1.baseline.md863 B
  • eval-1.with-skill.md2.4 KB
  • eval-2.baseline.md1.1 KB
  • eval-2.with-skill.md3.1 KB
  • eval-3.baseline.md632 B
  • eval-3.with-skill.md1.9 KB
  • eval-4.baseline.md667 B
  • eval-4.with-skill.md759 B
  • eval-5.baseline.md638 B
  • eval-5.with-skill.md1023 B
  • pressure-1.baseline.md676 B
  • pressure-1.with-skill.md810 B
  • pressure-2.baseline.md536 B
  • pressure-2.with-skill.md630 B
  • trigger-1.md160 B
  • trigger-2.md180 B
  • trigger-3.md162 B
  • trigger-4.md156 B
  • trigger-5.md140 B

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…