Skip to content
Back to skills

Proteomics Data Import

ASecurity

Load when ingesting a MaxQuant `proteinGroups.txt`, FragPipe `combined_protein.tsv`, DIA-NN

  • 8 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 6, 2026
datapythongobash

Works with

  • cli

Security analysis

A100/100

Pro scans all 6 files and shows the line behind each finding

Scanned September 6, 2026

npx -y skills add lilinji/GeneTind-Life-Skills --skill proteomics-data-import --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Proteomics Data Import?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Proteomics Data Import
[![Security: A β€” Skills Directory](https://www.skillsdirectory.com/api/skills/lilinji-proteomics-data-import/badge)](https://www.skillsdirectory.com/skills/lilinji-proteomics-data-import)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
# AUTO-GENERATED header from skill.yaml β€” do not edit by hand.
# Edit skill.yaml, then run: python scripts/generate_skill_md.py <skill_dir>
name: proteomics-data-import
description: Load when ingesting a MaxQuant `proteinGroups.txt`, FragPipe `combined_protein.tsv`, DIA-NN
  report, or generic CSV / TSV protein-quantification table β€” normalises columns to a standard schema,
  emits `tables/proteins.csv`. Skip when raw spectra are the input (run the search engine first); the
  file is already OmicsClaw schema.
version: 0.5.0
author: OmicsClaw
license: MIT
emoji: πŸ“₯
tags:
- proteomics
- import
- maxquant
- fragpipe
- diann
- spectronaut
requires:
- numpy
- pandas
---

# proteomics-data-import

## When to use

The user has a search-engine output (MaxQuant `proteinGroups.txt`,
FragPipe `combined_protein.tsv`, DIA-NN main report, or a generic
CSV / TSV protein table) and wants it normalised into OmicsClaw's
standard schema (lowercase `protein_id` plus `LFQ_<sample>` /
`Int_<sample>` intensity columns derived from MaxQuant's
`LFQ intensity ...` / `Intensity ...` headers).
Pick the format with `--format {maxquant,fragpipe,diann,generic}`
(default `maxquant`).

For raw MS spectra (mzML / RAW), run a search engine first
(MaxQuant / FragPipe / DIA-NN) and feed THIS skill the resulting
table.

## Inputs & Outputs

<!-- AUTO-GENERATED from skill.yaml (interface) β€” do not edit by hand. Regenerate: python scripts/generate_skill_md.py <skill_dir> -->

**Inputs**

- File types: `.txt`, `.tsv`, `.csv`

**Outputs**

- `tables/proteins.csv`
- `report.md`
- `result.json`

## Flow

1. Load input (`--input <file>`) or generate a demo MaxQuant-shaped file (`--demo`).
2. Dispatch to the format-specific importer (`proteomics_data_import.py:164-174` `_dispatch_import`); supported keys are `maxquant`, `fragpipe`, `diann`, `generic`.
3. Rename columns: `LFQ intensity <sample>` β†’ `LFQ_<sample>` and `Intensity <sample>` β†’ `Int_<sample>` (`proteomics_data_import.py:85`); `Majority protein IDs` β†’ `protein_id`; `Gene names` β†’ `gene_name`; etc.
4. Write `tables/proteins.csv` (`proteomics_data_import.py:284`) + `report.md` + `result.json` (`:299`).

## Gotchas

- **`--format` value must match `_dispatch_import` keys exactly.** `proteomics_data_import.py:166-171` registers `maxquant`, `fragpipe`, `diann`, `generic`. An unknown value raises `ValueError("Unsupported format: ... Supported: ['maxquant', 'fragpipe', 'diann', 'generic']")` at `:173`. There is no `spectronaut` importer despite the legacy SKILL.md mention β€” use `--format generic` for Spectronaut and rename columns yourself.
- **`--input` REQUIRED unless `--demo`.** `proteomics_data_import.py:275` raises `ValueError("--input required when not using --demo")`. Non-existent paths raise `FileNotFoundError` from `pd.read_csv`.
- **Output schema is LOWERCASE.** Column renaming targets `protein_id`, `intensity_<sample>`, `gene_name` etc. Downstream skills (`proteomics-quantification`, `proteomics-de`) assume this casing. Verify after import with `head tables/proteins.csv`.
- **No deduplication of contaminants / decoys.** Contaminant (`CON_*`) and decoy (`REV_*`) rows are passed through unchanged. Filter them upstream with the search engine's `--keep-contaminants false` flag, or add a downstream `df = df[~df["protein_id"].str.startswith(("CON_", "REV_"))]` step.

## Key CLI

```bash
# Demo (synthetic MaxQuant-style)
python omicsclaw.py run proteomics-data-import --demo --output /tmp/import_demo

# Real MaxQuant output
python omicsclaw.py run proteomics-data-import \
  --input proteinGroups.txt --output results/ --format maxquant

# FragPipe combined_protein
python omicsclaw.py run proteomics-data-import \
  --input combined_protein.tsv --output results/ --format fragpipe

# DIA-NN main report
python omicsclaw.py run proteomics-data-import \
  --input report.tsv --output results/ --format diann

# Generic / Spectronaut (rename columns yourself first)
python omicsclaw.py run proteomics-data-import \
  --input my_table.csv --output results/ --format generic
```

## See also

- `references/parameters.md` β€” every CLI flag
- `references/methodology.md` β€” per-format column-mapping rules
- `references/output_contract.md` β€” `tables/proteins.csv` schema
- Adjacent skills: `proteomics-ms-qc` (downstream β€” QC the imported table), `proteomics-quantification` (downstream β€” compute LFQ / iBAQ / spectral count), `proteomics-identification` (parallel β€” peptide-level summary), `proteomics-de` (downstream β€” differential abundance after import)

Files in this skill

  • SKILL.md4.5 KB
  • proteomics_data_import.py10.6 KB
  • references/methodology.md869 B
  • references/output_contract.md648 B
  • references/parameters.md263 B
  • skill.yaml1 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…