Skip to content
Back to skills

Run Private Document Extraction Pipelines With Text Extract Api

ASecurity

Use Text Extract API when an agent needs to turn PDFs, Office files, or images into Markdown or structured JSON with local OCR, optional Ollama models, PII removal, and queued batch processing.

  • 36 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 2, 2026
ai-agentspythongofastapidockergitapibackendsecuritydocumentation

Works with

  • cli
  • api

Security analysis

A92/100
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro shows the line behind each finding and how to fix it

Scanned September 2, 2026

npx -y skills add agentskillexchange/skills --skill run-private-document-extraction-pipelines-with-text-extract-api --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Run Private Document Extraction Pipelines With Text Extract Api?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Run Private Document Extraction Pipelines With Text Extract Api
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/agentskillexchange-run-private-document-extraction-pipelines-with-tex/badge)](https://www.skillsdirectory.com/skills/agentskillexchange-run-private-document-extraction-pipelines-with-tex)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: "Run private document extraction pipelines with Text Extract API"
slug: "run-private-document-extraction-pipelines-with-text-extract-api"
description: "Use Text Extract API when an agent needs to turn PDFs, Office files, or images into Markdown or structured JSON with local OCR, optional Ollama models, PII removal, and queued batch processing."
github_stars: 3111
verification: "security_reviewed"
source: "https://github.com/CatchTheTornado/text-extract-api"
author: "CatchTheTornado"
publisher_type: "open_source"
category: "Data Extraction & Transformation"
framework: "Multi-Framework"
tool_ecosystem:
  github_repo: "CatchTheTornado/text-extract-api"
  github_stars: 3111
---

# Run private document extraction pipelines with Text Extract API

Use Text Extract API when an agent needs to turn PDFs, Office files, or images into Markdown or structured JSON with local OCR, optional Ollama models, PII removal, and queued batch processing.

## Prerequisites

Python, FastAPI, Celery, Redis, Docker, Ollama, OCR backend

## Installation

Use the upstream install or setup path that matches your environment:
- [Download and install Docker](https://www.docker.com/products/docker-desktop/)
- git clone https://github.com/CatchTheTornado/text-extract-api.git
- pip install -e .
- brew update && brew install libmagic poppler pkg-config ghostscript ffmpeg automake autoconf

Requirements and caveats from upstream:
- **No Cloud/external dependencies** all you need: PyTorch based OCR (EasyOCR) + Ollama are shipped and configured via docker-compose no data is sent outside your dev/server environment,
- python client/cli.py ocr_upload --file examples/example-mri.pdf --prompt_file examples/example-mri-2-json-prompt.txt
- python client/cli.py ocr_upload --file examples/example-invoice.pdf --prompt_file examples/example-invoice-remove-pii.txt

Basic usage or getting-started notes:
- Before running the example see [getting started](#getting-started)
- ![Converting MRI report to Markdown](./screenshots/example-1.png)
- ![Converting Invoice to JSON](./screenshots/example-2.png)

- Source: https://github.com/CatchTheTornado/text-extract-api
- Extracted from upstream docs: https://raw.githubusercontent.com/CatchTheTornado/text-extract-api/HEAD/README.md

## Documentation

- https://github.com/CatchTheTornado/text-extract-api

## Source

- [Agent Skill Exchange](https://agentskillexchange.com/skills/run-private-document-extraction-pipelines-with-text-extract-api/)

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…