Skip to content
Back to skills

Dataset Split Review

ASecurity

Audit the methodology used to split data into train, validation, and test sets.

  • 6 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added May 29, 2026
ai-agentsperformance

Security analysis

A100/100

Scanned May 29, 2026

npx -y skills add yeaight7/agent-powerups --skill dataset-split-review --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Dataset Split Review?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Dataset Split Review
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/yeaight7-dataset-split-review/badge)](https://www.skillsdirectory.com/skills/yeaight7-dataset-split-review)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: dataset-split-review
description: "Audit the methodology used to split data into train, validation, and test sets."
---

# Dataset Split Review

A random split is often the wrong split. Incorrect splitting causes massive overestimation of model performance.

## Review Protocol

1. **Time-Series Data**: If the data has a time component, `train_test_split` is strictly forbidden. You must use a chronological split to prevent the model from learning the future.
2. **Group Leakage**: If the dataset has multiple rows for a single user/patient/session, a standard split will put rows from the same user in both train and test. You must use GroupKFold or group-based splitting.
3. **Stratification**: For imbalanced datasets, verify that stratification is used to maintain the target distribution across all splits.
4. **Action**: Review the splitting code and explicitly verify Time, Group, and Stratification safety.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…