Skip to content
Back to skills

Biomed Clinical Data Ml And Bioinformatics

ASecurity

Use when building models or pipelines on clinical or genomic data: clinical data structures and standards, the technicals of clinical machine learning — the prevalence problem, calibration rather than discrimination alone, where clinical models learn the wrong thing, and the validation ladder — and bioinformatics including sequence alignment, variant calling, and expression and single-cell analysis.

  • 2 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 19, 2026
ai-agentsgoexpressdatabaseperformance

Works with

  • cli

Security analysis

A100/100

Scanned September 19, 2026

npx -y skills add the-vibey-project/vibey --skill biomed-clinical-data-ml-and-bioinformatics --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Biomed Clinical Data Ml And Bioinformatics?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Biomed Clinical Data Ml And Bioinformatics
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/the-vibey-project-biomed-clinical-data-ml-and-bioinformatics/badge)](https://www.skillsdirectory.com/skills/the-vibey-project-biomed-clinical-data-ml-and-bioinformatics)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: biomed-clinical-data-ml-and-bioinformatics
description: "Use when building models or pipelines on clinical or genomic data: clinical data structures and standards, the technicals of clinical machine learning — the prevalence problem, calibration rather than discrimination alone, where clinical models learn the wrong thing, and the validation ladder — and bioinformatics including sequence alignment, variant calling, and expression and single-cell analysis."
---

# Biomedical Engineering: Clinical Data, Clinical ML, and Bioinformatics

> **Part 2 of 5** of the *Biomedical Engineering* reference (plugin `biomedical-engineering-technical`), covering §3–§5. Sibling skills: `biomed-signals-and-medical-imaging` (§0–§2), `biomed-structural-systems-biology-and-pharmacology` (§6–§9), `biomed-biomechanics-devices-and-biostatistics` (§10–§15), `biomed-reference` (§16–§20). Section numbers are shared across the set; a reference written as §N → `skill` points into that sibling skill.
>
> **Currency:** The physics, physiology and mathematics are stable; tool and pipeline recommendations shift slowly.

> **Scope note.** This is the engineering and science. **Regulatory pathways, quality
> systems, and lifecycle process are deliberately excluded** — they're a separate subject
> and they'd swamp the technical content.
>
> **⚠️ GOTCHA** boxes mark where a silent wrong answer is produced — which in this domain
> is the dangerous failure mode, not a crash.
>
> **The three technical facts that recur everywhere below:**
> 1. **⚠️ Biological signals are non-stationary, and most DSP assumes stationarity.** Every
>    windowing choice is an assumption about how long the physiology holds still (§1 → `biomed-signals-and-medical-imaging`).
> 2. **⚠️ Prevalence governs predictive value.** Sensitivity and specificity are properties
>    of a test; PPV is a property of a test *in a population*. Confusing them is the single
>    most common quantitative error in the field (§4.1).
> 3. **⚠️ Biological variability is the signal's competitor.** Between-subject variance
>    usually exceeds the effect you're measuring, which is why normalization,
>    within-subject designs, and mixed-effects models dominate (§15 → `biomed-biomechanics-devices-and-biostatistics`).

---

## §3. Clinical Data Structures

**[Technical view only — what the formats are and where they break.]**

| Standard | Structure |
|---|---|
| **HL7 v2** | ⚠️ **Pipe-delimited segments (MSH, PID, OBR, OBX). 1989. Still runs most hospitals** |
| **FHIR** | REST + JSON/XML, resource-oriented (Patient, Observation, Condition…) |
| **DICOM** | §2.2 → `biomed-signals-and-medical-imaging` |
| **OMOP CDM** | ⚠️ **Relational common data model for observational research — person, visit_occurrence, condition_occurrence, measurement, drug_exposure** |
| **CDISC SDTM/ADaM** | Trial submission tabulations |
| **SNOMED CT** | ⚠️ **~350k concepts, polyhierarchy, post-coordination** |
| **LOINC** | Lab/observation codes — 6-part name (component, property, time, system, scale, method) |
| **RxNorm** | Drug normalization across brand/generic/ingredient |
| **UCUM** | ⚠️ **Unit codes. Use them.** |

**⚠️ The technical failure modes:**
- **"FHIR-compliant" constrains almost nothing without a profile** (US Core, IPS,
  national). Two conformant servers can be mutually unusable.
- **⚠️ Terminology mapping is lossy and permanent.** A local lab code mapped to the wrong
  LOINC produces confidently wrong data downstream forever, and nothing will flag it.
- **⚠️ Units.** Glucose in mg/dL vs mmol/L differs by ~18×. Creatinine, calcium, and
  haemoglobin all have dual-unit conventions. **Store units with values, always.**
- **Patient identity is probabilistic** in systems without a universal identifier.
  ⚠️ **Both duplicates and wrong-patient merges occur.**
- **⚠️ Negation and hedging.** "No evidence of pneumonia" contains "pneumonia."
  **NegEx/ConText-style algorithms or a model that handles scope — naive keyword matching
  inverts meaning.**
- **Partial dates** ("1954") and timezone-naive timestamps are common and valid.
- **⚠️ EHR data is generated by billing and workflow, not by measurement.** The presence of
  a lab result is itself informative (§4.3) and rarely missing at random.

---

## §4. Clinical ML Technicals

### 4.1 ⚠️ The prevalence problem

**Sensitivity and specificity are properties of a test. PPV is a property of a test in a
population.**
```
PPV = (Sens × Prev) / (Sens × Prev + (1−Spec) × (1−Prev))
```
**Worked — a genuinely good test in a screening setting:**
```
Sens 95%, Spec 95%, Prevalence 1%
  PPV = (0.95 × 0.01) / (0.95×0.01 + 0.05×0.99)
      = 0.0095 / (0.0095 + 0.0495) = 16.1%     ⚠️
```
**⚠️ 84% of positives are false, with a 95/95 test.** Move prevalence to 20% (a referral
population) and PPV becomes 82.6%. **Same test. Nothing changed but the population.**

**⚠️ Consequences**: **accuracy is meaningless at low prevalence** (99% by always saying
no); **AUROC is prevalence-independent and therefore hides this**; **use AUPRC for
imbalanced problems**, whose baseline is the prevalence itself.

### 4.2 ⚠️ Calibration, not just discrimination

**Discrimination** (AUROC) asks whether ranking is correct. **Calibration** asks whether
"70%" means 70%. **For decision support, calibration is what the clinician acts on.**
- **Calibration curve** (predicted vs observed in bins), **calibration slope and
  intercept**, **Brier score** `(1/N)Σ(p_i − y_i)²`.
- ⚠️ **Modern neural networks are systematically overconfident.** **Temperature scaling**,
  Platt scaling, or isotonic regression on a held-out set.
- **⚠️ Calibration does not transfer across sites**, because prevalence differs (§4.1).
  **Recalibration at deployment site is usually necessary.**

**Decision curve analysis** — net benefit across threshold probabilities. ⚠️ **The metric
that actually answers "is this useful," and it's underused.**

### 4.3 ⚠️ Where clinical models learn the wrong thing

- **Shortcut features.** Documented: models keying on **chest-drain markers to predict
  pneumothorax**, **rulers in dermatology images to predict melanoma**, and **scanner or
  hospital identity** rather than pathology.
- **⚠️ Label leakage from the future.** *"Was a troponin ordered"* predicts myocardial
  infarction — because the clinician already suspected it. **Any feature downstream of
  clinical suspicion leaks.**
- **⚠️ Missingness is informative.** A missing lab is a decision not to order it.
  **Imputing it as MAR destroys signal and introduces bias simultaneously.** Model
  missingness explicitly (indicator variables) or use methods that handle it natively.
- **Treatment in the features** — confounding by indication.
- **Proxy targets encode system bias.** ⚠️ **A widely-deployed care-management algorithm
  used healthcare *spending* as a proxy for health *need*, thereby encoding unequal access
  as lower need.** The general lesson: **the target variable is a modelling choice with
  consequences.**
- **⚠️ Distribution shift** — scanner, protocol, coding practice, population, season.

### 4.4 The validation ladder
```
Internal (CV, holdout)     ⚠️ necessary, never sufficient
Temporal (train past, test future)   catches drift and leakage
External (different sites/scanners)  ⚠️ where performance typically drops
Prospective (real workflow, real time)
Outcome trial              ⚠️ does it change outcomes, not just AUROC?
```
**⚠️ Group-aware splitting is mandatory**: split by *patient*, not by image or record.
**A patient with 40 slices in both train and test leaks.**

**Reporting**: TRIPOD+AI, CONSORT-AI, SPIRIT-AI, STARD-AI.

---

## §5. Bioinformatics

### 5.1 Sequence alignment

**Pairwise, dynamic programming**: **Needleman–Wunsch** (global, `O(mn)`),
**Smith–Waterman** (local — ⚠️ **the difference is a zero floor in the recurrence and
traceback from the maximum cell**), with **affine gap penalties** (`gap_open` +
`k·gap_extend`) via Gotoh's three-matrix formulation, because a single long indel is
biologically far more likely than many separate ones.

**Substitution matrices**: **BLOSUM62** (protein, from blocks of aligned segments —
⚠️ **the default for most protein searches**), PAM series, and simple match/mismatch for
DNA.

**Heuristic search**: **BLAST** — seed with exact word matches, extend, evaluate.
⚠️ **E-value = expected number of hits this good by chance in a database this size**, so
**it depends on database size** — the same alignment gets a different E-value in a bigger
database.

**Read alignment at scale**: **BWA-MEM**, **Bowtie2**, **minimap2** (⚠️ **long reads**) —
all built on the **FM-index / Burrows-Wheeler transform**, giving `O(m)` substring search
in a compressed index. **This data structure is why resequencing became tractable.**

### 5.2 Variant calling

```
FASTQ → QC/trim → align (BAM) → mark duplicates → base recalibration
      → call variants (VCF) → filter → annotate → interpret
```
**Callers**: GATK HaplotypeCaller (⚠️ **local reassembly rather than pure pileup**),
DeepVariant (CNN over pileup images), FreeBayes, Strelka2. **Structural variants**:
Manta, DELLY, and long reads, which are far better at them.

**⚠️ The technical difficulties:**
- **Repetitive regions and segmental duplications** — ⚠️ **mapping quality collapses; reads
  map ambiguously and variants there are unreliable.**
- **Indel realignment** around true indels, which otherwise generate false SNVs.
- **Strand bias, allele balance** — a heterozygote should be ~50%; deviation signals
  artifact. **Somatic calling has no such expectation** because of tumour purity and
  subclonality.
- **⚠️ Coverage depth requirements differ enormously**: ~30× for germline SNVs, **but
  100–1000× for somatic variants at low allele fraction.**
- **Reference bias** — reads carrying the alternate allele map slightly worse.
  ⚠️ **Pangenome graph references (e.g. GRCh38 → draft human pangenome) address this** and
  are the direction of travel.

**Formats**: FASTQ (⚠️ **Phred quality `Q = −10 log₁₀ P(error)`; Q30 = 1 in 1000**),
SAM/BAM/CRAM, VCF, BED, GFF/GTF.

### 5.3 Expression and single-cell
**Bulk RNA-seq**: quantify (Salmon/kallisto — ⚠️ **pseudoalignment, dramatically faster**),
then differential expression with **DESeq2** or **edgeR** — ⚠️ **which model counts as
negative binomial, because RNA-seq is overdispersed relative to Poisson.**
**⚠️ Normalization matters more than the test**: TPM within-sample, but **median-of-ratios
or TMM across samples** — RPKM/FPKM are not comparable between samples.

**Single-cell**: QC (⚠️ **MAD-based filtering on counts, genes, and mitochondrial
fraction**) → normalize → HVG selection → PCA → neighbourhood graph → **Leiden
clustering** → UMAP.
> **⚠️ GOTCHA — UMAP and t-SNE distances are not meaningful.** Cluster *sizes*,
> inter-cluster *distances*, and apparent density are artifacts of the embedding.
> **Use them to visualize, never to quantify.** Conclusions must come from the graph or
> the expression, not the picture.

**Batch effects** are pervasive: Harmony, scVI, Combat. ⚠️ **Over-correction merges real
biology; under-correction leaves batch as the dominant axis. There is no automatic
answer.**

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…