Skip to content
Back to skills

Ai Jailbreak System Prompts

ASecurity

Advanced techniques for bypassing LLM safety filters, instruction tuning, and system prompt restrictions using specialized linguistic constructs, hypothetical scenarios, and persona adoption.

  • 22 stars
  • 0 votes
  • 0 copies
  • 3 views
  • Added September 12, 2026
ai-agentspythongotestingbackendsecurity

Security analysis

A100/100

Pro scans all 3 files and shows the line behind each finding

Scanned September 12, 2026

npx -y skills add ShulkwiSEC/bb-huge --skill ai-jailbreak-system-prompts --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Ai Jailbreak System Prompts?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Ai Jailbreak System Prompts
[![Security: A β€” Skills Directory](https://www.skillsdirectory.com/api/skills/shulkwisec-ai-jailbreak-system-prompts/badge)](https://www.skillsdirectory.com/skills/shulkwisec-ai-jailbreak-system-prompts)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: ai-jailbreak-system-prompts
description: >
  Advanced techniques for bypassing LLM safety filters, instruction tuning, and system prompt 
  restrictions using specialized linguistic constructs, hypothetical scenarios, and persona adoption.
domain: cybersecurity
subdomain: ai-red-teaming
category: Prompt Engineering
difficulty: advanced
estimated_time: "2 hours"
mitre_atlas:
  tactics: [AML.TA0001]
  techniques: [AML.T0043, AML.T0051]
mitre_attack:
  tactics: [TA0009]
  techniques: [T1592]
platforms: [ai, web]
tags: [ai, jailbreak, prompt-engineering, llm, safety-bypasses, red-teaming]
tools: [chat-interfaces, burp-suite, custom-scripts]
version: "1.0"
author: CyberSkills-Elite
license: Apache-2.0
---

# AI Jailbreaking & System Prompt Bypasses

## When to Use
- When conducting security assessments of Large Language Models (LLMs) integrated into chatbots, virtual assistants, or backend AI data processing pipelines.
- To demonstrate how instruction-tuned models can be forced into producing harmful, unethical, or restricted outputs by carefully crafting adversarial prompts.


## Prerequisites
- Access to target AI/ML system or local model deployment for testing
- Python 3.9+ with relevant ML libraries (transformers, torch, openai)
- Understanding of LLM architecture and prompt processing pipelines
- Authorized scope and rules of engagement for AI red team testing

## Workflow

### Phase 1: Understanding Target Model Constraints

```text
# Concept: LLM safety filters ```

### Phase 2: Persona Adoption Attacks

```text
# ```

### Phase 3: Developer Mode & Fictional Scenarios

```text
# ```

### Phase 4: Payload Encoding & Obfuscation

```text
# ```

#### Decision Point πŸ”€
```mermaid
flowchart TD
    A[Craft Prompt ] --> B{Bypass Successful ]}
    B -->|Yes| C[Capture Output ]
    B -->|No| D[Refine Approach ]
    C --> E[Test Edge Cases ]
```

## πŸ”΅ Blue Team Detection & Defense
- **Filter Ensembling**: **Context Monitoring**: Key Concepts
| Concept | Description |
|---------|-------------|
## Output Format
```
Ai Jailbreak System Prompts β€” Assessment Report
============================================================
Target: [Target identifier]
Assessor: [Operator name]
Date: [Assessment date]
Scope: [Authorized scope]
MITRE ATT&CK: [Relevant technique IDs]

Findings Summary:
  [Finding 1]: [Severity] β€” [Brief description]
  [Finding 2]: [Severity] β€” [Brief description]

Detailed Results:
  Phase 1: [Phase name]
    - Result: [Outcome]
    - Evidence: [Screenshot/log reference]
    - Impact: [Business impact assessment]

  Phase 2: [Phase name]
    - Result: [Outcome]
    - Evidence: [Screenshot/log reference]
    - Impact: [Business impact assessment]

Risk Rating: [Critical/High/Medium/Low/Informational]
Recommendations:
  1. [Immediate remediation step]
  2. [Long-term hardening measure]
  3. [Monitoring/detection improvement]
```


## πŸ“š Shared Resources
> For cross-cutting methodology applicable to all vulnerability classes, see:
> - [`_shared/references/elite-chaining-strategy.md`](../_shared/references/elite-chaining-strategy.md) β€” Exploit chaining methodology and high-payout chain patterns
> - [`_shared/references/elite-report-writing.md`](../_shared/references/elite-report-writing.md) β€” HackerOne-optimized report writing, CWE quick reference
> - [`_shared/references/real-world-bounties.md`](../_shared/references/real-world-bounties.md) β€” Verified disclosed bounties by vulnerability class

## References
- OWASP: [LLM Top 10 - Prompt Injection](https://owasp.org/www-project-machine-learning-security-top-10/)
- Anthropic: [Red Teaming Language Models](https://www.anthropic.com/index/red-teaming-language-models)

Files in this skill

  • SKILL.md3.6 KB
  • evals/evals.json540 B
  • scripts/process.py7.8 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…