Skip to content
Back to skills

Document Chunker

ASecurity

Split documents into overlapping token chunks for RAG pipelines and LLM context windows. Zero dependencies.

  • 6 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 4, 2026
ai-agentspythongobashgit

Works with

  • cli

Security analysis

A100/100

Pro scans all 13 files and shows the line behind each finding

Scanned September 4, 2026

npx -y skills add ellmos-ai/skills --skill document-chunker --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Document Chunker?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Document Chunker
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/ellmos-ai-document-chunker/badge)](https://www.skillsdirectory.com/skills/ellmos-ai-document-chunker)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: document-chunker
version: 1.0.0
type: tool
author: Lukas Geiger
created: 2026-03-12
updated: 2026-03-12
description: Split documents into overlapping token chunks for RAG pipelines and LLM context windows. Zero dependencies.

standalone: true
anthropic_compatible: true
bach_compatible: true
bach_origin: true
category: utilities
tags: [chunking, rag, tokens, nlp, text-processing, embedding]
language: de
status: active
visibility: public
dependencies: {'tools': [], 'services': [], 'protocols': [], 'python': []}
provenance: {'origin': 'bach', 'origin_path': 'system/tools/document_chunker.py', 'origin_version': '1.0.0', 'origin_repo': 'github.com/ellmos-ai/bach', 'last_sync_from_origin': '2026-03-12', 'last_sync_to_origin': None, 'local_changes_since_sync': False}
---

<img src="banner.png" width="100%" alt="document-chunker banner">

> **Deutsch** — Offizielle Deutsch-Version / Documento Oficial en Deutsch.


# Document Chunker (Deutsch)

Splits documents into overlapping token chunks. Optimized for RAG pipelines
and LLM context windows. Zero dependencies — Python stdlib + re only.

## Usage

### As Library
```python
from document_chunker import DocumentChunker

chunker = DocumentChunker(chunk_size=400, overlap=80)
chunks = chunker.chunk_text("Long text...")

for chunk in chunks:
    print(f"Chunk {chunk['chunk_id']}: {chunk['tokens']} tokens")
```

### Chunking a File
```python
chunks = chunker.chunk_document("document.md", source="My Project")
```

### Chunking an Entire Directory
```python
from document_chunker import chunk_corpus

chunks = chunk_corpus(["doc1.md", "doc2.txt"], source="Corpus")
```

### CLI
```bash
python document_chunker.py document.md    # Single file
python document_chunker.py ./docs/        # Entire directory
```

## Parameters

| Parameter | Default | Description |
|-----------|---------|-------------|
| chunk_size | 400 | Max tokens per chunk |
| overlap | 80 | Overlapping tokens between chunks |

## Supported File Types

`.txt`, `.md`, `.py`, `.sh`

## Änderungsprotokoll

### 1.0.0 (2026-03-12)
- Ported from BACH system/tools/document_chunker.py

Files in this skill

  • SKILL.en.md2 KB
  • SKILL.es.md2.2 KB
  • SKILL.fr.md2.2 KB
  • SKILL.ja.md2.4 KB
  • SKILL.md2.1 KB
  • SKILL.ru.md2.7 KB
  • SKILL.zh.md2.1 KB
  • document_chunker.py5.3 KB
  • en/SKILL.md2 KB
  • es/SKILL.md2.1 KB
  • ja/SKILL.md2.3 KB
  • ru/SKILL.md2.6 KB
  • zh/SKILL.md2 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…