Skip to content
Back to skills

Safety Scan

ASecurity

Scan inputs for prompt injection, unsafe content, and adversarial attacks using AIDefence

  • 73,733 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added May 27, 2026
data-airustgobashapi

Works with

  • api
  • mcp

Security analysis

A100/100

Scanned May 27, 2026

npx -y skills add ruvnet/ruflo --skill safety-scan --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Safety Scan?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Safety Scan
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/ruvnet-safety-scan/badge)](https://www.skillsdirectory.com/skills/ruvnet-safety-scan)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: safety-scan
description: Scan inputs for prompt injection, unsafe content, and adversarial attacks using AIDefence
argument-hint: "<input-text>"
allowed-tools: mcp__claude-flow__aidefence_scan mcp__claude-flow__aidefence_analyze mcp__claude-flow__aidefence_is_safe mcp__claude-flow__aidefence_learn mcp__claude-flow__aidefence_stats Bash
---

# Safety Scan

Scan content for prompt injection, jailbreak attempts, and unsafe patterns.

## When to use

Before processing untrusted input (user submissions, API payloads, webhook data), scan it to detect prompt injection, adversarial content, or policy violations.

## Steps

1. **Quick safety check** — call `mcp__claude-flow__aidefence_is_safe` with the input text for a boolean safe/unsafe result
2. **Deep analysis** — call `mcp__claude-flow__aidefence_analyze` for detailed threat classification and confidence scores
3. **Full scan** — call `mcp__claude-flow__aidefence_scan` for comprehensive multi-layer scanning
4. **Train defenses** — call `mcp__claude-flow__aidefence_learn` with confirmed threats to improve detection
5. **View stats** — call `mcp__claude-flow__aidefence_stats` for detection rates and false positive metrics

## Threat categories

- Prompt injection (direct and indirect)
- Jailbreak attempts
- Data exfiltration patterns
- Instruction override attacks
- Social engineering prompts

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…