Skip to content
Back to skills

Proteomics Enrichment

ASecurity

Load when running over-representation analysis (ORA) on a list of proteins via Fisher's exact

  • 8 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 6, 2026
datapythongobashreactdatabasebackend

Works with

  • cli

Security analysis

A100/100

Pro scans all 6 files and shows the line behind each finding

Scanned September 6, 2026

npx -y skills add lilinji/GeneTind-Life-Skills --skill proteomics-enrichment --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Proteomics Enrichment?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Proteomics Enrichment
[![Security: A β€” Skills Directory](https://www.skillsdirectory.com/api/skills/lilinji-proteomics-enrichment/badge)](https://www.skillsdirectory.com/skills/lilinji-proteomics-enrichment)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
# AUTO-GENERATED header from skill.yaml β€” do not edit by hand.
# Edit skill.yaml, then run: python scripts/generate_skill_md.py <skill_dir>
name: proteomics-enrichment
description: Load when running over-representation analysis (ORA) on a list of proteins via Fisher's exact
  test against a built-in 8-pathway DEMO dictionary, with BH-FDR correction. Skip when needing a real
  pathway database (this skill is demo-only) (use bulkrna-enrichment); rank-based GSEA.
version: 0.5.0
author: OmicsClaw
license: MIT
emoji: πŸ—ΊοΈ
tags:
- proteomics
- enrichment
- ora
- fisher
- demo
- pathway
requires:
- numpy
- pandas
- scipy
---

# proteomics-enrichment

## When to use

The user has a CSV listing proteins of interest (e.g. the
significant subset from `proteomics-de`, or the PTM-target list
from `proteomics-ptm`) and wants over-representation enrichment
via Fisher's exact test, with BH-adjusted FDR.

**This is a demo-only enrichment.** The pathway database is the
hard-coded 8-pathway `DEMO_PATHWAYS` dict at
`prot_enrichment.py:40-49` (each pathway has 5 fixed members).
There is NO CLI flag to load a real KEGG / Reactome / MSigDB
library. For production proteomics enrichment, export your
significant-protein list and call `bulkrna-enrichment` (which has
real ORA + GSEA + ssGSEA backends with hosted libraries).

## Inputs & Outputs

<!-- AUTO-GENERATED from skill.yaml (interface) β€” do not edit by hand. Regenerate: python scripts/generate_skill_md.py <skill_dir> -->

**Inputs**

- File types: `.csv`
- Accepts artifact `proteomics.differential_results` (`csv`)

**Outputs**

- `tables/enrichment_results.csv`
- `report.md`
- `result.json`

## Flow

1. Load CSV (`--input <proteins.csv>`) or generate a demo (`--demo`).
2. Pick the gene-list column: `protein_id` if present, otherwise the first column (`prot_enrichment.py:256`).
3. For each pathway in `DEMO_PATHWAYS` (`prot_enrichment.py:40-49`), run Fisher's exact test (`prot_enrichment.py:138-150`); apply BH FDR adjustment (`prot_enrichment.py:167`).
4. Write `tables/enrichment_results.csv` (`prot_enrichment.py:267`) + `report.md` + `result.json` (`:277`).

## Gotchas

- **Pathway database is HARD-CODED 8 demo pathways.** `prot_enrichment.py:40-49` defines 8 pathways Γ— 5 genes each (e.g. cell-cycle, apoptosis, TCA-cycle). There is no CLI for loading real databases. The `n_pathways_tested = 8` in `result.json` (`:271`) is a constant, not a function of input. For real enrichment, route to `bulkrna-enrichment`.
- **Method is Fisher's exact, not hypergeometric.** Mathematically equivalent for over-representation, but the script and report (`prot_enrichment.py:4, 138-150`) consistently say "Fisher's". Hypergeometric is the same distribution but the "Fisher's exact test" naming is what shows in the report.
- **Default background β‰  a real proteome size.** `prot_enrichment.py:126-128` sets `background_size = max(len(gene_set | all_pathway_genes), len(gene_set) + 1)` β€” for the demo's 8 pathways that's ~40 + n_input. **Always pass `--background-size N` (e.g. 20000 for human, 8000 for your detected proteome) for real enrichment** β€” the auto-default produces meaningless p-values on a real dataset.
- **`--species` is RECORDED-ONLY.** `prot_enrichment.py:237-241` accepts `--species` but the value is never used to switch databases or filter pathways β€” it's logged into `result.json` for reproducibility only.
- **Gene-list column auto-detection: `protein_id` first, else first column.** `prot_enrichment.py:256` uses `gene_col = "protein_id" if "protein_id" in df.columns else df.columns[0]`. If your CSV has multiple ID columns (`gene`, `uniprot`, `symbol`), only `protein_id` is preferred β€” pre-rename the column you want enriched.
- **`--input` REQUIRED unless `--demo`.** `prot_enrichment.py:251` raises `ValueError("--input required when not using --demo")`.

## Key CLI

```bash
# Demo (8-pathway DEMO_PATHWAYS dict)
python omicsclaw.py run proteomics-enrichment --demo --output /tmp/enr_demo

# Real protein list β€” pass real background size!
python omicsclaw.py run proteomics-enrichment \
  --input significant.csv --output results/ \
  --background-size 20000

# For real pathway databases, use bulkrna-enrichment instead:
# python omicsclaw.py run bulkrna-enrichment --input significant.csv ...
```

## See also

- `references/parameters.md` β€” every CLI flag
- `references/methodology.md` β€” Fisher's exact ORA, BH FDR, demo-DB caveats
- `references/output_contract.md` β€” `tables/enrichment_results.csv` schema
- Adjacent skills: `proteomics-de` (upstream β€” produces significant protein lists), `proteomics-ptm` (upstream β€” PTM-target lists), `proteomics-quantification` (upstream β€” protein-level abundance), `bulkrna-enrichment` (parallel β€” REAL pathway databases + GSEA + ssGSEA)

Files in this skill

  • SKILL.md4.7 KB
  • prot_enrichment.py9.9 KB
  • references/methodology.md836 B
  • references/output_contract.md649 B
  • references/parameters.md299 B
  • skill.yaml1.1 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…