Skip to content
Back to skills

Bulkrna De

ASecurity

Load when comparing gene expression between two conditions in bulk RNA-seq count data. Skip

  • 8 stars
  • 0 votes
  • 0 copies
  • 3 views
  • Added September 6, 2026
datapythongobashexpresstesting

Works with

  • cli

Security analysis

A100/100

Pro scans all 7 files and shows the line behind each finding

Scanned September 6, 2026

npx -y skills add lilinji/GeneTind-Life-Skills --skill bulkrna-de --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Bulkrna De?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Bulkrna De
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/lilinji-bulkrna-de/badge)](https://www.skillsdirectory.com/skills/lilinji-bulkrna-de)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
# AUTO-GENERATED header from skill.yaml — do not edit by hand.
# Edit skill.yaml, then run: python scripts/generate_skill_md.py <skill_dir>
name: bulkrna-de
description: Load when comparing gene expression between two conditions in bulk RNA-seq count data. Skip
  when the data is single-cell (use sc-de); spatial (use spatial-de); you need exon-level alternative
  splicing (use bulkrna-splicing).
version: 0.3.0
author: OmicsClaw
license: MIT
emoji: 🔬
tags:
- bulkrna
- differential-expression
- DESeq2
- volcano
- MA-plot
- fold-change
requires:
- matplotlib
- numpy
- pandas
- scipy
---

# bulkrna-de

## When to use

The user has a bulk RNA-seq count matrix (genes × samples) and wants to know
which genes change between two groups (control vs treatment, tumour vs normal,
etc.).  PyDESeq2 is preferred when ≥2 replicates per condition exist; Welch's
t-test is the fallback for the single-replicate / no-PyDESeq2 case.

## Inputs & Outputs

<!-- AUTO-GENERATED from skill.yaml (interface) — do not edit by hand. Regenerate: python scripts/generate_skill_md.py <skill_dir> -->

**Inputs**

- File types: `.csv`
- Accepts artifact `bulkrna.count_matrix` (`csv`)
- Tabular structure: at least 3 columns

**Outputs**

- `tables/counts.csv`
- `tables/de_results.csv`
- `tables/de_significant.csv`
- `tables/deseq2_results.csv`
- `figures/de_barplot.png`
- `figures/ma_plot.png`
- `figures/pvalue_histogram.png`
- `figures/volcano_plot.png`
- `report.md`
- `result.json`
- Produces artifact `bulkrna.differential_results` as `tables/de_results.csv` (`csv`)

## Flow

1. Load genes × samples raw count matrix.
2. Auto-partition columns into control / treatment by name prefix.
3. Pre-filter genes with total counts < 10 across all samples.
4. Run PyDESeq2 (negative binomial GLM + Wald test) — fall back to Welch's t-test if PyDESeq2 missing or < 2 replicates per condition.
5. Apply Benjamini–Hochberg FDR correction.
6. Filter DEGs by `--padj-cutoff` and `--lfc-cutoff`.
7. Render volcano / MA / p-value histogram and emit report.

## Gotchas

- **PyDESeq2 silently falls back to Welch's t-test** when fewer than 2 replicates per condition are detected or `pydeseq2` is not importable.  Check `result.json["method_used"]` to confirm which engine actually ran — the volcano-plot title alone does not surface the fallback.
- **LFCs are unshrunk by design.** Suitable for hypothesis testing (padj thresholds), but for ranking / visualisation that emphasises high-confidence effects, apply apeglm or ashr shrinkage *outside* this skill.
- **VST / rlog transformations are visualisation-only.** Do not feed transformed counts back into this skill — DE testing always wants raw integer counts.
- **Sample group detection is prefix-based.** Columns must start with `--control-prefix` (default `ctrl`) or `--treat-prefix` (default `treat`); columns matching neither prefix are silently dropped.
- **Pre-filter removes low-count genes** (total < 10).  This improves dispersion estimation but means the input gene count is not the testing gene count — `result.json["n_tested"]` is authoritative.

## Key CLI

```bash
# Demo run (synthetic 200-gene × 12-sample dataset)
python omicsclaw.py run bulkrna-de --demo

# Realistic run with custom prefixes and stricter cutoffs
python omicsclaw.py run bulkrna-de \
  --input counts.csv --output results/ \
  --control-prefix wt --treat-prefix ko \
  --padj-cutoff 0.01 --lfc-cutoff 1.5
```

## See also

- `references/parameters.md` — every CLI flag and tuning hint
- `references/methodology.md` — PyDESeq2 vs t-test, design validation, LFC shrinkage and transformation guidance
- `references/output_contract.md` — exact output directory layout
- Adjacent skills: `bulkrna-qc` (upstream count-matrix QC), `bulkrna-enrichment` (downstream pathway enrichment of DEG lists), `bulkrna-coexpression` (parallel WGCNA), `bulkrna-splicing` (exon-level alternative splicing)

Files in this skill

  • SKILL.md3.8 KB
  • bulkrna_de.py18.1 KB
  • references/methodology.md4.2 KB
  • references/output_contract.md1.6 KB
  • references/parameters.md338 B
  • skill.yaml1.4 KB
  • tests/test_bulkrna_de.py1.8 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…