Skip to content
Back to skills

Experiment Tracking Review

ASecurity

Use when reviewing ML training code to confirm a run could be reconstructed later -- hyperparameters, metrics, data references, and artifacts logged, not just printed to stdout.

  • 6 stars
  • 0 votes
  • 0 copies
  • 3 views
  • Added September 6, 2026
ai-agents

Security analysis

A100/100

Scanned September 6, 2026

npx -y skills add yeaight7/agent-powerups --skill experiment-tracking-review --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Experiment Tracking Review?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Experiment Tracking Review
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/yeaight7-experiment-tracking-review-agent-powerups/badge)](https://www.skillsdirectory.com/skills/yeaight7-experiment-tracking-review-agent-powerups)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: experiment-tracking-review
description: Use when reviewing ML training code to confirm a run could be reconstructed later -- hyperparameters, metrics, data references, and artifacts logged, not just printed to stdout.
---

## Purpose

An ML experiment is useless if you cannot reconstruct exactly how it was run and what data it used. This review checks that everything needed for reconstruction is logged.

## When to Use

- Reviewing a training script before serious experiment cycles begin
- Results exist but nobody can say which configuration produced them
- Metrics are printed to stdout only

## Inputs

- The training script(s) and any tracking/config setup (MLflow, wandb, config files)

## Workflow

1. **Hyperparameter logging**: the script logs *every* hyperparameter (learning rate, batch size, architecture details). Hardcoded magic numbers must be extracted to a config and logged.
2. **Metric logging**: training and validation metrics are logged at each epoch or step, not just at the end.
3. **Artifact saving**: final model weights, preprocessing scalers/encoders, and the exact configuration file are saved together in a versioned directory or tracking system.
4. **Data reference**: the run records which data it used (path, version, or hash).
5. **Enforce structured logging** (JSON, MLflow, wandb) — stdout-only metric reporting is a blocking finding.

## Output

- A finding list of what is not logged or saved, with the specific fix for each gap

## Verification

- [ ] Every hyperparameter logged (no unlogged magic numbers remain)
- [ ] Per-epoch/step train and validation metrics logged
- [ ] Weights + preprocessors + config saved together, versioned
- [ ] Data reference (path/version/hash) recorded with the run
- [ ] No stdout-only metric reporting remains

## Failure Modes

- **End-only metrics** — a final score without curves hides divergence and the onset of overfitting.
- **Orphaned artifacts** — weights saved without the matching scaler/encoder and config cannot be served correctly later.
- **Config drift** — the logged config differs from what the script actually used; log the resolved config at runtime.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…