Skip to content
Back to skills

Coverage Stratified Split

ASecurity

Stratify train/validation split by binned mask coverage percentage to ensure balanced foreground representation in segmentation tasks

  • 61 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 12, 2026
testingpython

Security analysis

A100/100

Scanned September 12, 2026

npx -y skills add wenmin-wu/ds-skills --skill coverage-stratified-split --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Coverage Stratified Split?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Coverage Stratified Split
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/wenmin-wu-coverage-stratified-split/badge)](https://www.skillsdirectory.com/skills/wenmin-wu-coverage-stratified-split)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: cv-coverage-stratified-split
description: Stratify train/validation split by binned mask coverage percentage to ensure balanced foreground representation in segmentation tasks
---

# Coverage-Stratified Split

## Overview

In segmentation tasks, naive random splits can produce folds with unbalanced foreground/background ratios — some folds get mostly empty masks, others get mostly full masks. Compute per-image mask coverage (foreground pixel ratio), bin into discrete classes, and use stratified splitting on these bins. This ensures each fold sees the full range of mask densities.

## Quick Start

```python
import numpy as np
from sklearn.model_selection import train_test_split

coverage = masks.sum(axis=(1, 2)) / (masks.shape[1] * masks.shape[2])

def coverage_to_class(val):
    for i in range(0, 11):
        if val * 10 <= i:
            return i
    return 10

coverage_classes = np.array([coverage_to_class(c) for c in coverage])

X_train, X_val, y_train, y_val = train_test_split(
    images, masks, test_size=0.2,
    stratify=coverage_classes, random_state=42
)
```

## Workflow

1. Compute mask coverage ratio for each training image (sum of foreground pixels / total pixels)
2. Bin coverage into discrete classes (e.g., 0-10% → class 0, 10-20% → class 1, ...)
3. Use binned classes as `stratify` parameter in `train_test_split` or `StratifiedKFold`
4. Validate that each fold has similar coverage distribution

## Key Decisions

- **10 bins**: covers 0-100% in 10% increments — fine enough for most tasks
- **Empty mask handling**: images with 0% coverage form their own bin, preventing empty-mask imbalance
- **vs random split**: critical when dataset has skewed coverage distribution (many empty masks)
- **With KFold**: use `StratifiedKFold(n_splits=5).split(X, coverage_classes)` for cross-validation

## References

- [U-net, dropout, augmentation, stratification](https://www.kaggle.com/code/phoenigs/u-net-dropout-augmentation-stratification)
- [U-net with simple ResNet Blocks v2 (New loss)](https://www.kaggle.com/code/shaojiaxin/u-net-with-simple-resnet-blocks-v2-new-loss)

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…