Skip to content
Back to skills

Literature

ASecurity

Load when extracting GEO accessions, dataset metadata, and downloadable references from a

  • 8 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 6, 2026
datapythongobash

Works with

  • cli

Security analysis

A100/100

Pro scans all 12 files and shows the line behind each finding

Scanned September 6, 2026

npx -y skills add lilinji/GeneTind-Life-Skills --skill literature --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Literature?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Literature
[![Security: A β€” Skills Directory](https://www.skillsdirectory.com/api/skills/lilinji-literature/badge)](https://www.skillsdirectory.com/skills/lilinji-literature)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
# AUTO-GENERATED header from skill.yaml β€” do not edit by hand.
# Edit skill.yaml, then run: python scripts/generate_skill_md.py <skill_dir>
name: literature
description: Load when extracting GEO accessions, dataset metadata, and downloadable references from a
  scientific paper (PDF / URL / DOI / PubMed ID / raw text) for downstream omics analysis. Skip when the
  dataset is already in hand; only routing a query (use orchestrator).
version: 0.5.0
author: OmicsClaw
license: MIT
emoji: πŸ“„
tags:
- literature
- pdf
- doi
- pubmed
- geo
- metadata
requires:
- requests
---

# literature

## When to use

The user provides a scientific paper reference (PDF path, URL,
DOI, PubMed ID, or raw text excerpt) and wants OmicsClaw to
extract GEO accessions, dataset metadata, and (optionally)
download referenced GEO datasets β€” so a downstream analysis skill
can be invoked on real data.

`--input-type` defaults to `auto` (sniffs from input shape).
`--no-download` skips the GEO download step (metadata only).

For dispatching a NL query to an analysis skill use `orchestrator`.
For scaffolding a new skill from a paper use `omics-skill-builder`.

## Inputs & Outputs

<!-- AUTO-GENERATED from skill.yaml (interface) β€” do not edit by hand. Regenerate: python scripts/generate_skill_md.py <skill_dir> -->

**Inputs**

- Input kinds: `file`, `freeform`
- File types: `.pdf`

**Outputs**

- `extracted_metadata.json`
- `report.md`
- `result.json`
- `<--data-dir>/<GSEid>/...`
- Produces artifact `literature.dataset_handoff` as `extracted_metadata.json` (`json`)

## Flow

1. Parse `--input` (or `--demo`); raise `parser.error('the following arguments are required: --input (unless --demo is used)')` at `literature_parse.py:38` when missing.
2. Detect input type (URL / DOI / PubMed / PDF / text) via `--input-type auto` or honour the explicit value.
3. Call `parse_input` (`skills/literature/core/parser.py`); fetch / parse content.
4. Call `extract_metadata` (`skills/literature/core/extractor.py`) β†’ identify GEO accessions, dataset metadata, study type.
5. If GEO accessions found AND not `--no-download`: call `download_geo_dataset` (`skills/literature/core/downloader.py`) β†’ save to `--data-dir`.
6. Write `extracted_metadata.json` (`literature_parse.py:80`) + `report.md` (`:193`) + `result.json` (`:147`).

## Gotchas

- **`--input` REQUIRED unless `--demo` β€” uses `parser.error` (exit 2).** `literature_parse.py:38` calls `parser.error('the following arguments are required: --input (unless --demo is used)')`. Different from most file-pipeline skills which raise `ValueError`.
- **`--input-type auto` heuristics are positional, not URL-aware.** `core/parser.py:35-55` checks the bare-DOI regex `^10\.\d{4,}/\S+` first; URLs always hit the `startswith("http")` branch and resolve to `url`, even when they wrap a DOI (`https://doi.org/10.1038/...`). For PDF / file paths use `--input-type file` explicitly β€” `Path.exists()` has to succeed for auto-detection to pick `file`.
- **GEO download requires internet access.** `download_geo_dataset` issues HTTP requests to GEO FTP. Air-gapped runs must pass `--no-download` or the run will hang / time out.
- **PDF parsing requires `pypdf` / similar.** If the PDF parser dependency is missing, the run errors out β€” verify `skills/literature/requirements.txt` is satisfied.
- **`extracted_metadata.json` is at `output_dir/` ROOT, not `tables/`.** This skill does NOT follow the `tables/<file>.csv` convention used by analysis skills.
- **Empty / unparseable input β‡’ exit 1 (not 2).** `literature_parse.py:64` calls `sys.exit(1)` on internal parse failure (distinct from the `parser.error` exit-2 path for missing args).

## Key CLI

```bash
# Demo (built-in local text)
python omicsclaw.py run literature --demo --output /tmp/lit_demo

# DOI
python omicsclaw.py run literature \
  --input "10.1038/s41586-021-03689-7" --output results/

# PDF (use --input-type file)
python omicsclaw.py run literature \
  --input my_paper.pdf --input-type file --output results/

# URL, metadata-only (no GEO download)
python omicsclaw.py run literature \
  --input "https://www.nature.com/articles/..." \
  --output results/ --no-download
```

## See also

- `references/parameters.md` β€” every CLI flag, input-type heuristics
- `references/methodology.md` β€” GEO accession rules, parser fallbacks
- `references/output_contract.md` β€” `extracted_metadata.json` schema
- Adjacent skills: `orchestrator` (downstream β€” routes the resulting dataset to an analysis skill), `omics-skill-builder` (parallel β€” scaffold a new skill from a paper)

Files in this skill

  • INDEX.md744 B
  • SKILL.md4.5 KB
  • core/__init__.py39 B
  • core/downloader.py5 KB
  • core/extractor.py10.7 KB
  • core/parser.py3.6 KB
  • literature_parse.py8.7 KB
  • references/methodology.md2.4 KB
  • references/output_contract.md1.2 KB
  • references/parameters.md300 B
  • requirements.txt53 B
  • skill.yaml1.3 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…