Skip to content
Back to skills

Transductive Train Test Transform

ASecurity

Fit unsupervised transforms (scaler, PCA, variance filter) on combined train+test data for more stable statistics, especially on small datasets

  • 61 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 12, 2026
datapythongo

Security analysis

A100/100

Scanned September 12, 2026

npx -y skills add wenmin-wu/ds-skills --skill transductive-train-test-transform --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Transductive Train Test Transform?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Transductive Train Test Transform
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/wenmin-wu-transductive-train-test-transform/badge)](https://www.skillsdirectory.com/skills/wenmin-wu-transductive-train-test-transform)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: tabular-transductive-train-test-transform
description: Fit unsupervised transforms (scaler, PCA, variance filter) on combined train+test data for more stable statistics, especially on small datasets
---

# Transductive Train-Test Transform

## Overview

When the training set is small, fitting a scaler or dimensionality reduction on train alone produces noisy statistics. For unsupervised transforms (StandardScaler, PCA, VarianceThreshold, NMF) that don't use the target, fitting on combined train+test is safe and produces more stable estimates. This is especially valuable when data is partitioned into many small subgroups.

## Quick Start

```python
import pandas as pd
from sklearn.preprocessing import StandardScaler
from sklearn.feature_selection import VarianceThreshold
from sklearn.pipeline import Pipeline

cols = [c for c in train.columns if c not in ['id', 'target', 'group']]

for g in train['group'].unique():
    tr = train[train['group'] == g][cols]
    te = test[test['group'] == g][cols]

    combined = pd.concat([tr, te])
    pipe = Pipeline([
        ('vt', VarianceThreshold(threshold=1.5)),
        ('scaler', StandardScaler())
    ])
    combined_t = pipe.fit_transform(combined)
    X_train = combined_t[:len(tr)]
    X_test = combined_t[len(tr):]
    # ... train model ...
```

## Workflow

1. Concatenate train and test feature matrices (exclude target and ID columns)
2. Fit unsupervised transforms on the combined data
3. Transform both portions
4. Split back into train and test by original lengths
5. Train supervised model on the transformed train set only

## Key Decisions

- **Safe transforms**: StandardScaler, MinMaxScaler, PCA, VarianceThreshold, NMF — all unsupervised, no leakage
- **Unsafe transforms**: target encoding, frequency encoding on target — DO NOT fit on test
- **When beneficial**: small datasets (< 1000 rows), partitioned data with < 100 rows per group
- **Large datasets**: minimal benefit — train-only fitting is already stable

## References

- [Pseudo Labeling - QDA - [0.969]](https://www.kaggle.com/code/cdeotte/pseudo-labeling-qda-0-969)
- [Quadratic Discriminant Analysis](https://www.kaggle.com/code/speedwagon/quadratic-discriminant-analysis)

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…