Skip to content
Back to skills

Ai Pentesting

ASecurity

Run autonomous AI-driven penetration tests on web applications using tools like Shannon, PentAGI, and similar frameworks. Use when tasks involve setting up automated penetration testing pipelines, combining AI agents with security tools (nmap, subfinder, nuclei, sqlmap), building autonomous exploit chains, generating pentest reports with proof-of-concept exploits, or integrating AI pentesting into CI/CD pipelines. Covers the full pentest lifecycle from reconnaissance to reporting using AI orc...

  • 142 stars
  • 0 votes
  • 0 copies
  • 8 views
  • Added May 27, 2026
testing-securitypythonrustgobashsqldockerawstestinggitapi

Works with

  • terminal
  • cli
  • api

Security analysis

A100/100

Pro scans all 2 files and shows the line behind each finding

Scanned October 4, 2026

npx -y skills add TerminalSkills/skills --skill ai-pentesting --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Ai Pentesting?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Ai Pentesting
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/terminalskills-ai-pentesting/badge)](https://www.skillsdirectory.com/skills/terminalskills-ai-pentesting)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: ai-pentesting
description: >-
  Run autonomous AI-driven penetration tests on web applications using tools
  like Shannon, PentAGI, and similar frameworks. Use when tasks involve setting
  up automated penetration testing pipelines, combining AI agents with security
  tools (nmap, subfinder, nuclei, sqlmap), building autonomous exploit chains,
  generating pentest reports with proof-of-concept exploits, or integrating
  AI pentesting into CI/CD pipelines. Covers the full pentest lifecycle from
  reconnaissance to reporting using AI orchestration.
license: Apache-2.0
compatibility: "Docker, Python 3.10+, Go 1.21+ for security tools"
metadata:
  author: terminal-skills
  version: "1.1.0"
  category: devops
  tags:
    - pentesting
    - ai-security
    - autonomous
    - vulnerability
    - owasp
---

# AI Pentesting

## Overview

Use AI agents to autonomously conduct penetration tests on web applications. Combine LLM reasoning with security tools (nmap, subfinder, nuclei, sqlmap, browser automation) to find and prove vulnerabilities with minimal human intervention.

## Instructions

### Methodology

AI pentesting follows the same phases as human pentesting, but the AI decides,
within each phase, which tools to run and what to try next based on findings:

1. **Reconnaissance** — subdomain enumeration (subfinder), technology
   fingerprinting (whatweb), port scanning (nmap), API schema discovery
   (crawling, OpenAPI/GraphQL introspection), and source-code analysis in
   white-box runs.
2. **Vulnerability analysis** — known-CVE scanning (nuclei), web scanning (OWASP
   ZAP, nikto), API fuzzing (schemathesis), and code-level hunting (semgrep,
   CodeQL) with input-to-sink data-flow analysis. The AI ranks what is likely
   exploitable.
3. **Exploitation** — proves candidates (injection, XSS, SSRF, auth bypass,
   business-logic flaws) using tools like sqlmap and browser automation
   (Playwright). The AI chooses order and chaining.
4. **Reporting** — a proof-of-concept per finding, reproducible steps, a CVSS
   severity, remediation guidance, and an executive summary.

### Setting Up Shannon

Shannon (`KeygraphHQ/shannon`, AGPL-3.0) is an autonomous AI pentester: it reads
your source code, identifies attack paths, and runs real exploits against a
running app to prove them. It is provider-agnostic (Anthropic, OpenAI, xAI,
Bedrock) and executes the run inside an ephemeral Docker worker. Run it only
against applications you own or have explicit written authorization to test —
not production.

```bash
# Configure AI-provider credentials with the interactive wizard (bring your own key)
npx @keygraph/shannon@latest setup

# -u target URL, -r source repo. The Docker worker cannot reach the host's localhost: use host.docker.internal
npx @keygraph/shannon@latest start \
  -u http://host.docker.internal:8080 \
  -r ./storefront-api

# Optional: supply a custom model configuration
#   npx @keygraph/shannon@latest start -u http://host.docker.internal:8080 -r ./storefront-api --models-config ./models.json
```

Shannon runs specialized agents for reconnaissance, per-category vulnerability
analysis (injection, XSS, SSRF, auth bypass), browser-based exploitation that
proves each finding with a real exploit, and reporting. Output includes vetted
findings with copy-paste PoC steps, a PDF report, and SARIF for code scanning.

### Building a Custom AI Pentest Pipeline

For cases where Shannon doesn't fit, build a custom pipeline:

```python
# ai_pentester.py
# Custom AI pentesting pipeline using LLM + security tools

import subprocess
import json
from openai import OpenAI

client = OpenAI()

class AIPentester:
    """Orchestrates security tools with LLM reasoning to find and prove bugs."""

    def __init__(self, target_url: str, scope: list[str] = None):
        self.target = target_url
        self.scope = scope or [target_url]
        self.findings = []
        self.recon_data = {}

    async def run_pentest(self) -> dict:
        """Execute the full lifecycle; returns findings, evidence, fixes."""
        # Phase 1: Recon
        self.recon_data = await self._recon()
        
        # Phase 2: AI-guided vulnerability analysis
        targets = await self._analyze_attack_surface(self.recon_data)
        
        # Phase 3: AI-guided exploitation
        for target in targets:
            finding = await self._exploit(target)
            if finding:
                self.findings.append(finding)
        
        # Phase 4: Generate report
        report = await self._generate_report()
        return report
    
    async def _recon(self) -> dict:
        """Run reconnaissance tools and aggregate results."""
        recon = {}
        
        # Subdomain enumeration
        result = subprocess.run(
            ['subfinder', '-d', self._get_domain(), '-silent'],
            capture_output=True, text=True, timeout=120
        )
        recon['subdomains'] = result.stdout.strip().split('\n')
        
        # Technology fingerprinting
        result = subprocess.run(
            ['whatweb', self.target, '--log-json=/dev/stdout', '-a', '3'],
            capture_output=True, text=True, timeout=60
        )
        recon['technologies'] = json.loads(result.stdout) if result.stdout else {}
        
        # Port scanning
        result = subprocess.run(
            ['nmap', '-sV', '--top-ports', '1000', '-oX', '-', self._get_domain()],
            capture_output=True, text=True, timeout=300
        )
        recon['ports'] = result.stdout
        
        # Nuclei scan for known CVEs
        # nuclei v3: use -jsonl (the old -json flag is deprecated)
        result = subprocess.run(
            ['nuclei', '-u', self.target, '-severity', 'critical,high',
             '-jsonl', '-silent'],
            capture_output=True, text=True, timeout=300
        )
        recon['known_vulns'] = [
            json.loads(line) for line in result.stdout.strip().split('\n')
            if line.strip()
        ]
        
        return recon
    
    async def _analyze_attack_surface(self, recon: dict) -> list:
        """Use AI to analyze recon data and prioritize attack targets."""
        response = client.chat.completions.create(
            model="gpt-4o",
            messages=[
                {"role": "system", "content":
                 "You are an expert penetration tester. Analyze the "
                 "reconnaissance data and identify the most promising "
                 "attack vectors. Return JSON array of targets."},
                {"role": "user", "content":
                 f"Recon data:\n{json.dumps(recon, indent=2)}\n\n"
                 "Identify attack targets with: endpoint, vulnerability_type, "
                 "technique, priority (1-5), reasoning."}
            ],
            response_format={"type": "json_object"}
        )
        return json.loads(response.choices[0].message.content).get("targets", [])

    async def _exploit(self, target: dict) -> dict | None:
        """Attempt to exploit an identified vulnerability."""
        vuln_type = target.get('vulnerability_type', '').lower()
        handlers = {
            'injection': self._test_injection,
            'xss': self._test_xss,
            'ssrf': self._test_ssrf,
            'auth': self._test_auth_bypass,
        }
        for key, handler in handlers.items():
            if key in vuln_type:
                return await handler(target)
        return None

    async def _generate_report(self) -> dict:
        """Generate a structured penetration test report."""
        response = client.chat.completions.create(
            model="gpt-4o",
            messages=[
                {"role": "system", "content":
                 "Generate a professional penetration test report with "
                 "executive summary, findings with CVSS scores, PoC steps, "
                 "and remediation recommendations."},
                {"role": "user", "content":
                 f"Target: {self.target}\n"
                 f"Findings: {json.dumps(self.findings, indent=2)}\n"
                 f"Recon data: {json.dumps(self.recon_data, indent=2)}"}
            ]
        )
        return {
            "target": self.target,
            "findings_count": len(self.findings),
            "findings": self.findings,
            "report": response.choices[0].message.content
        }
```

### CI/CD Integration

Run AI pentests on every deployment with the maintained `shannon-action`. Point
it at a staging instance you control, never production. The action refuses to
run in a public repository, because reports and logs contain findings.

```yaml
# .github/workflows/pentest.yml
name: AI Penetration Test
on:
  push:
    branches: [main]
  schedule:
    - cron: '0 2 * * 1'  # Weekly, Monday 02:00 UTC

permissions:
  contents: read
  security-events: write   # required for upload-sarif (code scanning)

jobs:
  pentest:
    runs-on: ubuntu-latest
    services:
      app:
        image: ghcr.io/myorg/storefront-api:${{ github.sha }}
        ports:
          - 8080:8080            # published on the runner host
    steps:
      - uses: actions/checkout@v4

      - name: Run Shannon
        uses: KeygraphHQ/shannon-action@v1
        with:
          url: http://host.docker.internal:8080   # repo defaults to the checkout
          api-key: ${{ secrets.SHANNON_AI_API_KEY }}
          fail-on-severity: high   # fail the build on high/critical findings
          upload-sarif: true       # publish findings to GitHub code scanning
```

### Report Structure

A professional AI-generated pentest report should include: executive summary (scope, duration, methodology, overall risk, findings count by severity), individual findings (each with CVSS score, affected endpoint/parameter, evidence with reproducible curl commands, impact description, and specific remediation guidance), and a remediation priority list ordered by severity with recommended fix timelines.

## Examples

### Run an autonomous pentest on a web application

```prompt
I run the storefront-api locally at http://localhost:8080 and the source is in
./storefront-api, which I own. Set up Shannon and run a full pentest, then
summarize the findings with reproducible proof-of-concept steps and flag any
critical issues.
```

The agent runs `npx @keygraph/shannon@latest setup`, then `start -u http://host.docker.internal:8080
-r ./storefront-api` (the worker container cannot reach the host's `localhost`),
and relays the vetted findings and PDF/SARIF report Shannon produces.

### Build a custom AI pentest pipeline

```prompt
Build a custom AI pentesting pipeline that combines subfinder, whatweb, nuclei
(with -jsonl output), and schemathesis, orchestrated by an LLM that analyzes
each tool's output and decides what to test next. Target a local app I own at
http://localhost:3000 with its OpenAPI spec at /docs/openapi.json. Produce a
structured findings report.
```

### Integrate AI pentesting into CI/CD

```prompt
Add automated pentesting to our GitHub Actions pipeline, running on push to main
and weekly. The app runs in Docker exposed at localhost:8080 on the runner. Use
the KeygraphHQ/shannon-action, publish findings to code scanning via SARIF, and
fail the build on high or critical findings.
```

## Guidelines

- Only run penetration tests against systems you have explicit written authorization to test — unauthorized testing is illegal
- AI pentesters can cause real damage (data modification, service disruption) — always test against staging environments, never production
- Review AI-generated exploitation attempts before running them — LLMs can hallucinate or generate overly aggressive payloads
- Treat pentest reports as confidential — they contain vulnerability details and proof-of-concept exploits
- Set time limits and scope boundaries for autonomous testing to prevent runaway scans
- Validate AI findings manually — false positives in automated reports erode trust with stakeholders
- Store API keys and credentials used for pentesting securely — never hardcode them in CI configurations

Files in this skill

  • SKILL.md12.3 KB
  • _scores.json1.6 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…