Skip to content
Back to skills

Anchor Grouped Validation

ASecurity

Split validation by unique anchor/query entities so no anchor appears in both train and val, preventing data leakage in pairwise matching tasks

  • 61 stars
  • 0 votes
  • 0 copies
  • 3 views
  • Added September 12, 2026
documentationpython

Security analysis

A100/100

Scanned September 12, 2026

npx -y skills add wenmin-wu/ds-skills --skill anchor-grouped-validation --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Anchor Grouped Validation?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Anchor Grouped Validation
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/wenmin-wu-anchor-grouped-validation/badge)](https://www.skillsdirectory.com/skills/wenmin-wu-anchor-grouped-validation)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: nlp-anchor-grouped-validation
description: Split validation by unique anchor/query entities so no anchor appears in both train and val, preventing data leakage in pairwise matching tasks
---

# Anchor-Grouped Validation

## Overview

In pairwise matching tasks (phrase similarity, question-answer matching), the same anchor/query maps to multiple targets. If the same anchor appears in both train and validation, the model memorizes anchor-specific patterns rather than learning general similarity. Splitting by unique anchors ensures the model is evaluated on truly unseen queries.

## Quick Start

```python
import numpy as np
from sklearn.model_selection import StratifiedGroupKFold

groups = df['anchor'].values
scores = (df['score'] * 100).astype(int)  # bin for stratification

sgkf = StratifiedGroupKFold(n_splits=4, shuffle=True, random_state=42)
for fold, (train_idx, val_idx) in enumerate(sgkf.split(df, scores, groups)):
    train_df = df.iloc[train_idx]
    val_df = df.iloc[val_idx]
    assert set(train_df['anchor']) & set(val_df['anchor']) == set()
```

## Workflow

1. Identify the grouping entity (anchor, query, user, document ID)
2. Use `StratifiedGroupKFold` for balanced + grouped splits, or manual shuffle-split on unique entities
3. Verify no group overlap between train and validation sets
4. Report both in-group (standard) and grouped CV scores — the gap reveals leakage risk

## Key Decisions

- **StratifiedGroupKFold vs GroupKFold**: stratified variant preserves label distribution across folds
- **Stratification target**: bin continuous scores into integers for stratification compatibility
- **Number of folds**: 4-5 is standard; fewer folds if number of unique anchors is small
- **Simple alternative**: shuffle unique anchors, split first 25% as val, rest as train

## References

- [Iterate like a grandmaster!](https://www.kaggle.com/code/jhoward/iterate-like-a-grandmaster)

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…