Skip to content
Back to skills

Proteomics Ms Qc

ASecurity

Load when computing protein-table QC — proteins × samples count, missing-value rate, intensity

  • 8 stars
  • 0 votes
  • 0 copies
  • 6 views
  • Added September 6, 2026
datapythongobash

Works with

  • cli

Security analysis

A100/100

Pro scans all 6 files and shows the line behind each finding

Scanned September 6, 2026

npx -y skills add lilinji/GeneTind-Life-Skills --skill proteomics-ms-qc --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Proteomics Ms Qc?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Proteomics Ms Qc
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/lilinji-proteomics-ms-qc/badge)](https://www.skillsdirectory.com/skills/lilinji-proteomics-ms-qc)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
# AUTO-GENERATED header from skill.yaml — do not edit by hand.
# Edit skill.yaml, then run: python scripts/generate_skill_md.py <skill_dir>
name: proteomics-ms-qc
description: Load when computing protein-table QC — proteins × samples count, missing-value rate, intensity
  CV (median + mean) — from a MaxQuant / FragPipe / DIA-NN protein-quantification CSV. Skip when raw mzML
  / RAW spectra are the input (run a search engine first); peptide-level QC is needed (use proteomics-identification).
version: 0.5.0
author: OmicsClaw
license: MIT
emoji: 📊
tags:
- proteomics
- qc
- ms
- maxquant
- intensity
- missing-values
- cv
requires:
- numpy
- pandas
---

# proteomics-ms-qc

## When to use

The user has a protein-quantification CSV (typically the output of
`proteomics-data-import`, with rows = proteins and columns =
samples + metadata) and wants QC summary statistics: protein count,
sample count, fraction of missing intensities, per-protein
coefficient of variation (CV) — median and mean. Auto-detects
intensity columns by `select_dtypes(include=[np.number])`.

This skill does NOT process raw spectra. For peptide / PSM-level
identification stats use `proteomics-identification`.

## Inputs & Outputs

<!-- AUTO-GENERATED from skill.yaml (interface) — do not edit by hand. Regenerate: python scripts/generate_skill_md.py <skill_dir> -->

**Inputs**

- Modalities: ms
- File types: `.csv`

**Outputs**

- `tables/qc_metrics.csv`
- `report.md`
- `result.json`

## Flow

1. Load CSV (`--input <file.csv>`) or generate a demo at `output_dir/demo_proteomics.csv` (`proteomics_ms_qc.py:223`).
2. Detect numeric (intensity) columns via `select_dtypes(include=[np.number])` (`proteomics_ms_qc.py:47`); raise `ValueError("No intensity/sample columns detected in input data")` at `:74` if none found.
3. Compute n_proteins / n_samples / missing_rate / per-protein CV.
4. Write `tables/qc_metrics.csv` (`proteomics_ms_qc.py:241`) + `report.md` + `result.json`.

## Gotchas

- **Sample columns must be NUMERIC.** Intensity-column auto-detection (`proteomics_ms_qc.py:47`) uses `select_dtypes(include=[np.number])`. String-typed intensities (e.g. quoted numbers in some Spectronaut exports) are silently treated as metadata, not samples — your `n_samples` will be 0 and the run raises `ValueError` at `:74`.
- **No intensity columns ⇒ hard fail.** `proteomics_ms_qc.py:74` raises `ValueError("No intensity/sample columns detected in input data")` — there is no auto-detection of `intensity_*` prefixes; only dtype-based.
- **`--input` REQUIRED unless `--demo`.** `proteomics_ms_qc.py:228` raises `ValueError("--input required when not using --demo")`.
- **Both `NaN` and `0.0` count as missing.** `proteomics_ms_qc.py:80` computes `missing_mask = np.isnan(intensities) | (intensities == 0)` — zero is treated as "not detected" (the proteomics convention). If your search engine writes a small placeholder (e.g. `1.0`) for undetected proteins, the missing rate is artificially LOW; pre-impute placeholders to `0` or `NaN` first.
- **CV is per-protein across samples.** Reported `median_cv` / `mean_cv` are aggregations across the per-protein CV distribution — interpret as "typical protein-level reproducibility", not "sample-level reproducibility".

## Key CLI

```bash
# Demo
python omicsclaw.py run proteomics-ms-qc --demo --output /tmp/qc_demo

# Real protein table (e.g. output of proteomics-data-import)
python omicsclaw.py run proteomics-ms-qc \
  --input results/tables/proteins.csv --output qc_results/
```

## See also

- `references/parameters.md` — every CLI flag
- `references/methodology.md` — CV definition, missing-value handling
- `references/output_contract.md` — `tables/qc_metrics.csv` schema
- Adjacent skills: `proteomics-data-import` (upstream — produces the protein table), `proteomics-quantification` (downstream — LFQ / iBAQ / spectral count), `proteomics-identification` (parallel — peptide-level summary), `proteomics-de` (downstream — differential abundance)

Files in this skill

  • SKILL.md3.9 KB
  • proteomics_ms_qc.py9.1 KB
  • references/methodology.md851 B
  • references/output_contract.md713 B
  • references/parameters.md326 B
  • skill.yaml1006 B

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…