Skip to content
Back to skills

Alterlab Pdf Extract

ASecurity

Free Elicit-columns analog — ingest N PDFs (or any MarkItDown-supported document) and build a per-paper evidence table with user-defined columns, one row per paper and one column per attribute/question you want pulled from every source. Use when extracting structured data across many papers into a comparison table or data-extraction sheet (sample size, methods, main finding, effect, population/intervention/outcome, limitations), screening a corpus into a spreadsheet, or pulling the same field...

  • 68 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 6, 2026
ai-agentspythongobashgitapidatabasebackend

Works with

  • cli
  • api

Security analysis

A92/100
  • mediumInstalls packages at runtime which could introduce malicious dependencies

Pro scans all 3 files and shows the line behind each finding

Scanned September 23, 2026

npx -y skills add AlterLab-IEU/AlterLab-Academic-Skills --skill alterlab-pdf-extract --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Alterlab Pdf Extract?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Alterlab Pdf Extract
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/alterlab-ieu-alterlab-pdf-extract/badge)](https://www.skillsdirectory.com/skills/alterlab-ieu-alterlab-pdf-extract)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: alterlab-pdf-extract
description: Free Elicit-columns analog — ingest N PDFs (or any MarkItDown-supported document) and build a per-paper evidence table with user-defined columns, one row per paper and one column per attribute/question you want pulled from every source. Use when extracting structured data across many papers into a comparison table or data-extraction sheet (sample size, methods, main finding, effect, population/intervention/outcome, limitations), screening a corpus into a spreadsheet, or pulling the same fields from a stack of PDFs into CSV/Markdown. Routes conversion through MarkItDown; offline heuristic backend by default, optional LLM backend for precise answers. Part of the AlterLab Academic Skills suite.
allowed-tools: Read Write Edit Bash
license: MIT
compatibility: markitdown >= 0.1 (uv pip install 'markitdown[all]') required for conversion; optional llm backend needs the openai client + OPENROUTER_API_KEY (or OPENAI_API_KEY with OPENAI_BASE_URL) and sends paper text to that endpoint
metadata:
    skill-author: AlterLab
    version: "1.0.1"
    last_updated: "2026-09-23"
---

# PDF Extract — User-Defined Evidence Tables

## Overview

The free analog of Elicit's "columns" feature. Point it at a stack of PDFs, declare the
columns you care about, and get back a tidy evidence table — **one row per paper, one
column per question** — as CSV or a Markdown table. Built for data extraction across a
corpus: literature screening, methods comparison, PICO/sample-size pulls, and review
evidence sheets.

Every input is normalized to clean Markdown via Microsoft **MarkItDown** (see the
`alterlab-markitdown` skill) before any column is extracted, so PDF, DOCX, PPTX, HTML, and
the other MarkItDown formats all work.

## When to Use This Skill

- Extracting the **same fields from many papers** into a comparison table or spreadsheet.
- Building a data-extraction sheet for a review (sample size, design, intervention,
  outcome, effect size, limitations).
- Screening a folder of PDFs into a CSV you can sort/filter.
- Any "make me a table where each row is a paper and each column is X" request.

This skill operates on documents you already have and produces a structured table, not
prose — use it *after* you have the papers.

### Does NOT Trigger

| Scenario | Use Instead |
|----------|-------------|
| Searching databases to find papers, or writing the narrative synthesis | `alterlab-literature-review` |
| Interrogating one PDF in depth (every section, figure, table, appendix) | `alterlab-pdf-explore` |
| Just converting a document to Markdown | `alterlab-markitdown` |
| Structured experimental data for papers you do not have on disk | `alterlab-bgpt-search` |
| Managing the PDFs and their references in a Zotero library | `alterlab-pyzotero` |

## Quick Start

```bash
SCRIPT=skills/research-tools/alterlab-pdf-extract/scripts/extract_to_table.py

# Heuristic backend: use 'label:question' so the question carries the words the
# papers actually use — overlap is lexical, so a bare label like "Sample size"
# rarely matches "we enrolled 240 participants" (see Backends below).
uv run python "$SCRIPT" papers/*.pdf \
    --column "Sample size:how many participants were enrolled" \
    --column "Main finding:primary result reported" \
    --column "Limitations:study limitations weaknesses" --format md

# Columns with explicit extraction questions, CSV to a file:
uv run python "$SCRIPT" a.pdf b.pdf \
    --column "n:how many participants were enrolled" \
    --column "effect:what was the main reported effect size" \
    --format csv -o evidence.csv

# Precise LLM-backed extraction (needs OPENROUTER_API_KEY) — label-only columns
# are fine here, since the model reads the paper rather than matching keywords:
uv run python "$SCRIPT" lit/*.pdf \
    --columns "Population" "Intervention" "Outcome" --backend llm
```

## Defining Columns

A column is a **label** plus the **question** that drives extraction. Two ways to declare:

| Spec form | Meaning |
|-----------|---------|
| `"Sample size"` | label == question; keywords come from the label |
| `"n:how many participants"` | label `n`, question `how many participants` |

Use `--columns` for a space-separated list of simple labels, and repeat `--column` for
specs that need an explicit question. Both combine; duplicate labels are rejected.

## Backends

- **heuristic** (default, offline, free): for each column, scores Markdown sentences by
  **literal keyword overlap** with the column spec and emits the best evidence snippet.
  Deterministic — good for a first pass, screening, and reproducible runs. Returns an empty
  cell when no sentence shares a keyword.
  **Gotcha:** matching is lexical, not semantic. A generic label like `"Sample size"`
  tokenizes to `sample`, `size` — which never appears in "we enrolled 240 participants", so
  the cell comes back empty. For the heuristic backend, phrase columns as `Label:question`
  where the question uses the **words the papers actually use** (`participants`, `enrolled`,
  `cohort`, `n=`). Quantitative questions also get a small ranking bonus for sentences
  containing digits when the question itself contains a digit or `%`. When you only have
  generic labels, use the `llm` backend instead.
- **llm** (optional, precise): sends the converted Markdown (first 24,000 characters) plus all
  column questions to an OpenAI-compatible endpoint for a concise per-column answer (`N/A`
  when the paper is silent). With `OPENROUTER_API_KEY` set, requests go to OpenRouter;
  otherwise `OPENAI_API_KEY` is sent to `OPENAI_BASE_URL` (or the OpenAI default) — each key
  only ever goes to its own endpoint. The model ID follows the **ALTERLAB_MODEL** convention
  (`skills/core/shared/model_env.md`): `$ALTERLAB_MODEL`, else the dated default; override
  per-run with `--model` (on a non-OpenRouter endpoint, pass a model ID it serves). Spot-check
  LLM answers against the PDFs before they enter a review, since a confident wrong number
  looks the same as a right one in the table.

## Output

- `--format md` → GitHub-flavored Markdown table (pipes in cell text are escaped).
- `--format csv` → CSV via `csv.DictWriter`.
- `-o FILE` writes to disk (UTF-8); otherwise the table goes to stdout.

The first column is always `Source` (the input filename); a file that fails conversion
yields an `ERROR` row so the table stays aligned.

## Pipeline Fit

`alterlab-pdf-extract` sits downstream of paper discovery and upstream of synthesis:

1. Find papers — database/search skills (`alterlab-literature-review`, search backends).
2. **Extract the evidence table — this skill.**
3. Synthesize / write up — `alterlab-literature-review`, `alterlab-paper-writer`.

## Files

- `scripts/extract_to_table.py` — the CLI. `--help` lists every flag; `--columns` /
  `--column` define the table; `--backend`, `--format`, `--model`, `-o` control extraction
  and output.

Part of the AlterLab Academic Skills suite.

Files in this skill

  • SKILL.md6 KB
  • evals/evals.json4.2 KB
  • scripts/extract_to_table.py13.6 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…