Skip to content
Back to skills

Bulkrna Qc

ASecurity

Load when checking a bulk RNA-seq count matrix for library-size outliers, gene detection

  • 8 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 6, 2026
datapythonrustgobashexpresstestinggit

Works with

  • cli

Security analysis

A100/100

Pro scans all 6 files and shows the line behind each finding

Scanned September 6, 2026

npx -y skills add lilinji/GeneTind-Life-Skills --skill bulkrna-qc --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Bulkrna Qc?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Bulkrna Qc
[![Security: A β€” Skills Directory](https://www.skillsdirectory.com/api/skills/lilinji-bulkrna-qc/badge)](https://www.skillsdirectory.com/skills/lilinji-bulkrna-qc)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
# AUTO-GENERATED header from skill.yaml β€” do not edit by hand.
# Edit skill.yaml, then run: python scripts/generate_skill_md.py <skill_dir>
name: bulkrna-qc
description: Load when checking a bulk RNA-seq count matrix for library-size outliers, gene detection
  rates, and sample-sample correlation before DE. Skip when data is raw FASTQ (use bulkrna-read-qc); aligner
  logs (use bulkrna-read-alignment); single-cell counts (use sc-qc).
version: 0.3.0
author: OmicsClaw
license: MIT
emoji: πŸ“Š
tags:
- bulkrna
- QC
- count-matrix
- library-size
- gene-detection
- sample-correlation
- CPM
requires:
- matplotlib
- numpy
- pandas
- scipy
---

# bulkrna-qc

## When to use

Run as the first step on a bulk RNA-seq count matrix (genes Γ— samples)
before differential expression.  Surfaces the four failure modes that
silently bias DE results: a sample with a tiny library, a sample with
suspiciously few detected genes, a low-correlation outlier vs the rest,
and CPM-vs-raw comparison artefacts.

## Inputs & Outputs

<!-- AUTO-GENERATED from skill.yaml (interface) β€” do not edit by hand. Regenerate: python scripts/generate_skill_md.py <skill_dir> -->

**Inputs**

- File types: `.csv`

**Outputs**

- `tables/cpm_normalized.csv`
- `tables/sample_stats.csv`
- `figures/expression_density.png`
- `figures/gene_detection.png`
- `figures/library_sizes.png`
- `figures/sample_correlation.png`
- `report.md`
- `result.json`

## Flow

1. Load the count matrix (raise on missing `--input` or non-existent file per `bulkrna_qc.py:428,431`).
2. Compute per-sample library sizes and detected-gene counts.
3. Compute sample Γ— sample correlation matrix; flag samples below the median-of-medians threshold as outliers.
4. Compute CPM normalisation as a side artifact (write `tables/cpm_normalized.csv`).
5. Render four figures and emit `report.md` + `result.json`.

## Gotchas

- **Hard-fails on missing input.**  `bulkrna_qc.py:428` raises `ValueError("--input is required when not using --demo")`; `:431` raises `FileNotFoundError` if the path doesn't exist.  No silent demo fallback when `--input` is given but invalid β€” fix the path or use `--demo`.
- **CPM is for visualisation only.**  `tables/cpm_normalized.csv` is emitted as a downstream-friendly artefact, but **DE testing must always use raw counts** (PyDESeq2's negative-binomial GLM expects integer counts; feeding CPM produces meaningless dispersion estimates).  Do not pipe `cpm_normalized.csv` into `bulkrna-de`.
- **Outlier flagging is correlation-based, not biology-aware.**  If two biological conditions differ strongly (e.g. tumour vs normal), the cross-condition correlations are *expected* to be lower β€” the outlier flag may fire on legitimate biology.  Cross-check `result.json["outlier_samples"]` against the experimental design before excluding samples.
- **First column is treated as the gene-id column unconditionally.**  If the CSV has a header row but no leading id column (samples-only), the first sample column will be silently parsed as gene names and omitted from QC.  Inspect `report.md`'s "samples seen" count vs your design before trusting the output.

## Key CLI

```bash
python omicsclaw.py run bulkrna-qc --demo
python omicsclaw.py run bulkrna-qc --input counts.csv --output results/
```

## See also

- `references/parameters.md` β€” every CLI flag and tuning hint
- `references/methodology.md` β€” library-size, gene-detection, correlation-based outlier metrics
- `references/output_contract.md` β€” exact output directory layout
- Adjacent skills: `bulkrna-read-qc` / `bulkrna-read-alignment` (upstream), `bulkrna-de` (downstream β€” raw counts only), `bulkrna-batch-correction` (downstream if QC reveals batch effects), `sc-qc` (single-cell sibling)

Files in this skill

  • SKILL.md3.7 KB
  • bulkrna_qc.py16.8 KB
  • references/methodology.md2.7 KB
  • references/output_contract.md1.4 KB
  • references/parameters.md326 B
  • skill.yaml1.2 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…