Skip to content
Back to skills

Reproducible Research

ASecurity

Make an analysis re-runnable by anyone: enforce notebook hygiene (top-to-bottom-clean, no out-of-order state), pin environments and dependencies in a lockfile, set and thread random seeds, version data and artifacts, track experiments (params + metrics + code/data version per run), and convert a run-once notebook into a scripted, deterministic pipeline.

  • 7 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 23, 2026
ai-agentspythonci/cd

Security analysis

A100/100

Scanned September 23, 2026

npx -y skills add mcorbett51090/RavenClaude --skill reproducible-research --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Reproducible Research?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Reproducible Research
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/mcorbett51090-reproducible-research/badge)](https://www.skillsdirectory.com/skills/mcorbett51090-reproducible-research)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: reproducible-research
description: "Make an analysis re-runnable by anyone: enforce notebook hygiene (top-to-bottom-clean, no out-of-order state), pin environments and dependencies in a lockfile, set and thread random seeds, version data and artifacts, track experiments (params + metrics + code/data version per run), and convert a run-once notebook into a scripted, deterministic pipeline."
---

# Reproducible Research

## Pin everything that can drift
Pin the Python version and every dependency in a lockfile (`requirements.txt` pinned / `poetry.lock` / conda env / container). A floating `>=` reproduces today and breaks on the next release. Deterministic install is the floor.

## Version the data, not just the name
"The data" must be a version, not a name. Snapshot or content-hash the exact input (DVC / immutable snapshot) so a result is recoverable; "the latest table" makes every past result unverifiable.

## Seed every stochastic step
A single global seed is not enough — the split, the model, the framework, and parallel ops each leak nondeterminism if you don't thread the seed through. Hunt the named sources of nondeterminism (unset seed, floating version, un-versioned data, thread-count-dependent ops) rather than shrugging at "flaky."

## Notebook is a draft; the pipeline is the deliverable
An exploratory notebook is scratch — out-of-order cells, hidden globals. Make it restart-and-run-all clean, or extract it to a scripted, deterministic pipeline. Don't confuse a run-once notebook with a result.

## Track the run, not just the result
Log params, metrics, artifacts, **and the code+data version** per run (MLflow or equivalent). Metrics without the commit and the data hash mean you can compare scores but never recover *why*.

## Output
A reproducibility deliverable: a pinned env, seeds set and threaded, versioned data, tracked runs, a clean notebook, and a scripted re-runnable pipeline. Hand production CI/CD and serving to `ml-engineering`; the analysis content stays with `exploratory-data-scientist` / `feature-and-modeling-engineer`.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…