Skip to content
Back to skills

Polars Dovmed

ASecurity

Run structured full-text searches of PMC Open Access and bioRxiv with polars-dovmed. Use for reproducible literature searches through the hosted API or local parquet corpora.

  • 9 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 27, 2026
ai-agentspythonshellbashgitapidatabase

Works with

  • api

Security analysis

A100/100

Pro scans all 6 files and shows the line behind each finding

Scanned October 2, 2026

npx -y skills add fmschulz/omics-skills --skill polars-dovmed --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Polars Dovmed?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Polars Dovmed
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/fmschulz-polars-dovmed/badge)](https://www.skillsdirectory.com/skills/fmschulz-polars-dovmed)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: polars-dovmed
description: Run structured full-text searches of PMC Open Access and bioRxiv with polars-dovmed. Use for reproducible literature searches through the hosted API or local parquet corpora.
user-invocable: true
---

# polars-dovmed

Search the PubMed Central Open Access and bioRxiv corpora with `polars-dovmed`, through the hosted API or local parquet files.

`SKILL_DIR` is the directory that contains this `SKILL.md`, installed at `~/.agents/skills/polars-dovmed`. Run the helper as `uv run --script "$SKILL_DIR/scripts/query_literature.py"` so its pinned dependency is available. The helper loads `~/.config/polars-dovmed/.env`.

## Access

- Hosted API: needs `POLARS_DOVMED_API_KEY` in the environment or in the `.env` file. Never pass the key in argv, and never write it to run directories, memory records, summaries, or answers. The API root (`https://api.newlineages.com/`) states that keys are issued by the admin only.
- Local corpora: upstream `dovmed download`, `dovmed build-parquet`, and `dovmed scan` from <https://github.com/UriNeri/polars-dovmed>. `omics-skills` ships neither an API key nor the corpora.
- With neither, state that `polars-dovmed` is not configured and use another literature skill.

## Instructions

1. Create a run directory, `tasks/polars-dovmed-runs/<date-topic>/`, outside `$SKILL_DIR`. Save `prompt.txt`, `query.json`, payloads, raw responses, timing files, and any curated summary there. Do not commit generated runs.
2. Write `query.json` (format below) and inspect it before searching: spelling, taxonomy aliases, regex escaping, and noisy terms.
   - Start anchor-only: exact entity names and evidence-based aliases. For prompts such as "hosts of X" or "distribution and genomics of X", search X and its aliases first and use the topic words for triage or a second pass. Never add generic words (distribution, genome, environment, host, taxonomy) as peer OR concepts.
   - Do not invent aliases by splitting compact taxon, gene, or clade names. Keep a real multi-word synonym as one phrase term.
   - Put acronym collisions and wrong systems in `disqualifying_terms`.
3. Run discovery first, one corpus per call. Never start with `--corpus both`.
   - bioRxiv: async discovery with `--poll-timeout 75` and `--details-rerank-limit` 8 to 12.
   - OpenPMC: `--sync --year-bands clean_split --year-band-workers 4 --skip-details-rerank --poll-timeout 120`. Map a requested year range to its bands and search only those (table below). Never run a broad unbanded OpenPMC scan. Fetch details by PMCID after inspecting the top hits.
   - Use `--extract-matches none --add-group-counts primary` for discovery.
   - Wrap calls in `timeout` when it exists.
4. Inspect the first 5 to 10 hits. If they are noisy, refine `query.json` and rerun.
5. Fill citation metadata from the corpus details first, then from `--crossref-metadata` or `/crossref-lookup`. Use DOI landing pages only when corpus metadata and Crossref disagree or stay ambiguous; do not use generic web search for DOI repair.
6. When speed matters, record wall time with shell `time -p` or `timeout`, plus the helper's `elapsed_ms`. Timeouts and failed endpoints are results; report them.
7. Do not use the flat `/api/search_literature` endpoint. It needs `--allow-flat-query` and can hang behind the edge proxy.

### Query JSON

```json
{
  "anchor_entity": [["primary_name"], ["alias_1"], ["alias_2"]],
  "relation_or_property": [["primary_name", "relation_term"]],
  "disqualifying_terms": [["term_to_exclude"]]
}
```

Both search paths (FTS and parquet scan) apply the same logic: terms in one inner group are AND'd and must co-occur; inner groups in a concept are OR'd; concepts are OR'd, and items matching more groups rank higher. A multi-word term matches as an exact phrase. List each synonym as its own single-term group; use a multi-term group only to require co-occurrence.

Rank hits in this order: anchor in title, anchor in abstract, anchor plus support term in title or abstract, several distinct group matches, full-text-only matches. Down-rank hits that match only generic support terms.

### Year bands (OpenPMC)

| Requested years | `--year-bands` value |
|---|---|
| 2009 and earlier | `pre_2010` |
| 2010 to 2020 | `2010_2020` |
| 2021 to 2023 | `2021_2023` |
| 2024 and later | `2024_plus` |
| No constraint | `clean_split` (all four) |

For a range that crosses bands, pass a comma list, for example `--year-bands 2021_2023,2024_plus`.

### Hosted API commands

```bash
RUN=tasks/polars-dovmed-runs/mirusviricota-hosts-$(date +%Y%m%d)
mkdir -p "$RUN"
printf '%s\n' "papers describing hosts of Mirusviricota" > "$RUN/prompt.txt"
# Write and inspect "$RUN/query.json" before searching.

# bioRxiv discovery
timeout 90s uv run --script "$SKILL_DIR/scripts/query_literature.py" \
  --queries-file "$RUN/query.json" --corpus biorxiv --mode discovery \
  --extract-matches none --add-group-counts primary \
  --max-results 25 --details-rerank-limit 12 --poll-timeout 75 \
  --save-payload "$RUN/payload_biorxiv.json" --save-response "$RUN/results_biorxiv.json" \
  > "$RUN/summary_biorxiv.json" 2> "$RUN/time_biorxiv.txt"

# OpenPMC discovery over the four clean year bands
timeout 120s uv run --script "$SKILL_DIR/scripts/query_literature.py" \
  --queries-file "$RUN/query.json" --corpus pmc --mode discovery \
  --sync --year-bands clean_split --year-band-workers 4 --skip-details-rerank \
  --extract-matches none --add-group-counts primary \
  --max-results 25 --poll-timeout 120 \
  --save-payload "$RUN/payload_pmc.json" --save-response "$RUN/results_pmc.json" \
  > "$RUN/summary_pmc.json" 2> "$RUN/time_pmc.txt"

# Details for known PMCIDs (bioRxiv: pass DOIs with --corpus biorxiv),
# with bounded Crossref fill-in for missing DOI, year, or journal
uv run --script "$SKILL_DIR/scripts/query_literature.py" \
  --details PMC6362216 PMC10132079 --corpus pmc \
  --crossref-metadata --crossref-limit 10 \
  --save-payload "$RUN/payload_details.json" --save-response "$RUN/results_details.json"
```

### Local commands

```bash
uv run --script "$SKILL_DIR/scripts/query_literature.py" \
  --execution-mode local --corpus pmc \
  --local-parquet-pattern "$DOVMED_PMC_PARQUET" \
  --queries-file "$RUN/query.json" \
  --save-payload "$RUN/payload_local.json" --save-response "$RUN/results_local.json"
```

- `--corpus pmc` reads `DOVMED_PMC_PARQUET` and `--corpus biorxiv` reads `DOVMED_BIORXIV_PARQUET` unless `--local-parquet-pattern` is given. `--corpus both` needs one compatible explicit pattern; otherwise run separate scans.
- Local runs record `--corpus-revision` (or `DOVMED_CORPUS_REVISION`) and read `flattened.csv`, falling back to `processed.parquet` (or the legacy `prcoessed.parquet`) only when the compact output is absent. `--timeout` caps a local scan (default 900 seconds).

### Reachability and smoke check

`curl -sS --max-time 20 https://api.newlineages.com/` returns service metadata. It proves the service is reachable, not that a corpus scan works. The smoke check runs one bounded bioRxiv discovery and one details lookup:

```bash
uv run --script "$SKILL_DIR/scripts/smoke_test.py" --run-dir tasks/polars-dovmed-runs/smoke-test
```

Run either check before declaring an outage.

## Input Requirements

- A literature question, turned into an inspected `query.json`.
- A hosted API key or local parquet files for the requested corpus.
- A writable run directory outside `$SKILL_DIR`.

## Output

- `prompt.txt`, `query.json`, payload JSON, raw response JSON, timing files, and an optional curated summary in the run directory.
- A paper list with titles, identifiers, corpus, relevance notes, and timing.
- Warnings for timeouts, HTTP 502, `Database not found`, missing metadata, or fallback use.

## Quality Gates

- [ ] The API key stayed in the environment and appears in no artifact.
- [ ] The run directory is outside `$SKILL_DIR`.
- [ ] `query.json` was written anchor-first and inspected before the search.
- [ ] Each corpus ran as its own discovery call; OpenPMC used clean year bands with `--sync` and `--skip-details-rerank`, limited to the requested years.
- [ ] The flat endpoint, `--corpus both`, and unbanded OpenPMC scans were avoided, or the reason is stated.
- [ ] Every hosted call had a poll timeout and, where available, a wall-clock `timeout`.
- [ ] The first 5 to 10 hits were reviewed and noisy terms refined.
- [ ] Citation metadata came from corpus details, then bounded Crossref, not generic web search.
- [ ] Timings and failures are reported when speed is part of the request.

## Troubleshooting

**Issue**: `--corpus both` returns HTTP 502.
**Solution**: Run `--corpus biorxiv` and `--corpus pmc` separately and merge the results.

**Issue**: OpenPMC year-band search returns `Database not found`.
**Solution**: Report that the deployment does not expose the indexed OpenPMC bands. Do not fall back to an unbanded scan.

**Issue**: OpenPMC year-band search times out.
**Solution**: Report the timeout and keep the bioRxiv result. Narrow to the requested bands or refine the anchors; do not retry with longer waits or an unbanded scan.

**Issue**: A "topic of X" query returns generic topic papers.
**Solution**: Rerun with the entity and its aliases only, and use topic terms during triage.

**Issue**: Raw `urllib` requests get HTTP 403 with Cloudflare error 1010.
**Solution**: Cloudflare rejects Python's default user agent. Use the helper, or send both `X-API-Key` and a `User-Agent` header.

**Issue**: `/usr/bin/time` is missing.
**Solution**: Use shell `time -p` or record start and end timestamps.

Files in this skill

  • SKILL.md18 KB
  • agents/openai.yaml301 B
  • fixtures/smoke_prompt.txt45 B
  • fixtures/smoke_query.json109 B
  • scripts/query_literature.py64.4 KB
  • scripts/smoke_test.py7.3 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…