Skip to content
Back to skills

Apache Tika Document Extractor

ASecurity

Wraps Apache Tika Server REST API for extracting structured text from PDFs, DOCX, PPTX, and 1,200+ file formats. Outputs clean markdown with metadata preservation using Tika /rmeta/text endpoint and recursive parsing mode.

  • 36 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added June 2, 2026
ai-agentsgojavadockergitapisecurity

Works with

  • api

Security analysis

A100/100

Scanned June 2, 2026

npx -y skills add agentskillexchange/skills --skill apache-tika-document-extractor --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Apache Tika Document Extractor?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Apache Tika Document Extractor
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/agentskillexchange-apache-tika-document-extractor/badge)](https://www.skillsdirectory.com/skills/agentskillexchange-apache-tika-document-extractor)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: "Apache Tika Document Extractor"
slug: "apache-tika-document-extractor"
description: "Wraps Apache Tika Server REST API for extracting structured text from PDFs, DOCX, PPTX, and 1,200+ file formats. Outputs clean markdown with metadata preservation using Tika /rmeta/text endpoint and recursive parsing mode."
github_stars: 3695
verification: "security_reviewed"
source: "https://github.com/apache/tika"
category: "Data Extraction & Transformation"
framework: "Codex"
tool_ecosystem:
  github_repo: "apache/tika"
  github_stars: 3695
---

# Apache Tika Document Extractor

Wraps Apache Tika Server REST API for extracting structured text from PDFs, DOCX, PPTX, and 1,200+ file formats. Outputs clean markdown with metadata preservation using Tika /rmeta/text endpoint and recursive parsing mode.

## Installation

Requirements and caveats from upstream:
- **N.B.** [Docker](https://www.docker.com/products/personal) is used for tests in tika-integration-tests. If Docker is not installed, those tests are skipped.

Basic usage or getting-started notes:
- ===========
- **Parse a file in Java:**
- java

- Source: https://github.com/apache/tika
- Extracted from upstream docs: https://raw.githubusercontent.com/apache/tika/HEAD/README.md

## Source

- [Agent Skill Exchange](https://agentskillexchange.com/skills/apache-tika-document-extractor/)

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…