Skip to content
Back to skills

Processing Pdfs

ASecurity

Processes PDF files. Extracts text and tables, fills forms, merges and splits documents, batch-processes files, converts to images, and generates PDFs programmatically. Use when working with .pdf files. Do NOT use for Word documents, spreadsheets, or presentations.

  • 239 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added May 26, 2026
ai-agentspythonbash

Works with

  • cli

Security analysis

A100/100

Pro scans all 13 files and shows the line behind each finding

Scanned May 27, 2026

npx -y skills add telagod/code-abyss --skill processing-pdfs --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Processing Pdfs?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Processing Pdfs
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/telagod-processing-pdfs/badge)](https://www.skillsdirectory.com/skills/telagod-processing-pdfs)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: processing-pdfs
description: Processes PDF files. Extracts text and tables, fills forms, merges and splits documents, batch-processes files, converts to images, and generates PDFs programmatically. Use when working with .pdf files. Do NOT use for Word documents, spreadsheets, or presentations.
user-invocable: false
allowed-tools: Bash, Read, Write, Edit, Glob
argument-hint: <file.pdf | task>
---

# PDF Processing

Essential PDF operations using Python libraries and CLI tools.

## Decision Matrix

| Task | Best Tool | Reference |
|------|-----------|-----------|
| Merge / split / metadata / rotate | pypdf | [recipes.md](references/recipes.md) |
| Extract text (layout preserved) | pdfplumber | [recipes.md](references/recipes.md) |
| Extract tables | pdfplumber | [recipes.md](references/recipes.md) |
| Create new PDF | reportlab | [recipes.md](references/recipes.md) |
| Batch CLI ops | qpdf / pdftk | [recipes.md](references/recipes.md) |
| OCR scanned PDFs | pytesseract + pdf2image | [advanced.md](references/advanced.md) |
| Add watermark / extract images / encrypt | pypdf / pdfimages | [advanced.md](references/advanced.md) |
| Fill PDF forms | pdf-lib / pypdf | [FORMS.md](FORMS.md) |
| Advanced pypdfium2 / pdf-lib JS | — | [REFERENCE.md](REFERENCE.md) |

## Quick Start

```python
from pypdf import PdfReader
reader = PdfReader("document.pdf")
print(f"Pages: {len(reader.pages)}")
text = "".join(page.extract_text() for page in reader.pages)
```

## Workflow

1. **Identify task** — text extraction? table? creation? form? Pick row from matrix above.
2. **Load reference** — recipes.md covers 90% of tasks; advanced.md for OCR / encrypt; FORMS.md for forms.
3. **Implement** — copy-adapt recipe; verify output.
4. **Validate** — open in a viewer or grep extracted text.

## Library Selection

| Library | Use for |
|---------|---------|
| pypdf | Merge, split, metadata, encryption, rotation |
| pdfplumber | Text extraction with layout, tables |
| reportlab | Generate PDFs programmatically |
| pdf2image + pytesseract | OCR scanned documents |
| qpdf / pdftk (CLI) | Batch ops, no Python needed |

Files in this skill

  • FORMS.md9.2 KB
  • REFERENCE.md16.3 KB
  • SKILL.md2.1 KB
  • references/advanced.md1.1 KB
  • references/recipes.md3.3 KB
  • scripts/check_bounding_boxes.py3.1 KB
  • scripts/check_bounding_boxes_test.py8.6 KB
  • scripts/check_fillable_fields.py362 B
  • scripts/convert_pdf_to_images.py1.1 KB
  • scripts/create_validation_image.py1.6 KB
  • scripts/extract_form_field_info.py6 KB
  • scripts/fill_fillable_fields.py4.7 KB
  • scripts/fill_pdf_form_with_annotations.py3.5 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…