Installs into .claude/skills of the current project.
Are you the author of Synthetic Data?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/aicodedecode-synthetic-data)
---
name: synthetic-data
description: Generate high-quality training data with models — pipelines, filtering, and validation for synthetic datasets that actually help.
category: ai-research
---
## Overview
Synthetic data is training data generated by models rather than collected from
humans or the world: instruction-response pairs, reasoning traces, code with
tests, preference judgments, tool-use trajectories. It's the engine behind most
modern post-training — distillation, self-improvement loops, and domain adaptation
all run on synthetic data. Done well, it converts a strong teacher model's
capabilities into targeted training signal for a student. Done poorly, it produces
fluent garbage that teaches the student the teacher's failure modes.
The core discipline is treating synthetic generation as a data pipeline with
quality gates, not as "prompt a big model and hope." Every synthetic dataset needs:
a generation step with controlled diversity, a verification step that filters for
correctness (execution, entailment checks, cross-model agreement), and a validation
step measuring downstream effect on the student. The failure mode to fear is model
collapse — training on unfiltered synthetic data degrades quality over
generations — which is why verification and mixing with real data matter.
Synthetic data is leverage: it multiplies a strong teacher into arbitrary
quantities of targeted training signal. Like all leverage, it amplifies errors
too. The pipeline's job is to keep the signal while discarding the errors.
## When to use
- Post-training a model for a domain where human data is scarce or expensive
(specialized reasoning, low-resource languages, niche tools).
- Distilling capabilities from a stronger teacher into a smaller, cheaper student.
- Generating instruction-tuning data at a scale humans can't match.
- Creating adversarial or edge-case examples to patch specific weaknesses found in
evals.
- Bootstrapping preference data (pairs judged by a strong model) for DPO or
reward modeling.
- Augmenting thin slices of real data (rare classes, edge cases) where collection
is impractical.
## Core concepts
- **Generation with diversity control.** Naive sampling collapses to the teacher's
modal outputs. Use temperature variation, persona/topic seeding, and explicit
diversity prompts ("generate a problem unlike these examples") to cover the
space.
- **Verification beats volume.** The highest-leverage step: filter generated data
for correctness. Executable domains (code, math) verify by running; open-ended
domains use entailment checks, multi-sample consistency, or a second model as
judge.
- **Rejection sampling / best-of-N.** Generate N candidates, keep the ones that
pass verification. This converts compute into quality and is the workhorse of
synthetic pipelines.
- **Distillation vs. self-improvement.** Distillation transfers from a stronger
teacher; self-improvement (STaR-style) has the model generate its own rationales
and keeps the ones leading to correct answers. Both need the same verification
discipline.
- **Model collapse.** Repeated training on unfiltered model outputs narrows the
distribution and amplifies errors. Mitigations: keep a substantial fraction of
real data, filter aggressively, and regenerate from the strongest available
teacher each round.
- **Contamination hygiene.** Synthetic eval-adjacent data must never leak into
benchmarks. Track provenance: which teacher, which prompt, which filter version
produced every example.
- **Teacher-student gap.** The student learns the teacher's verified outputs, not
the teacher's capabilities. A student trained on traces can match the teacher's
task performance without matching its generality — know which one you need.
- **Format diversity.** Varying output formats during generation prevents the
student from overfitting to one teacher's stylistic template.
## Practical workflow
1. **Define the target capability precisely.** "Better reasoning" is not a spec.
Write the task distribution: input types, difficulty mix, output format. The
generation prompts derive from this spec.
2. **Seed for diversity.** Build a seed pool of topics, personas, difficulty
levels, and edge cases. Sample seeds combinatorially so generation covers the
space rather than clustering.
3. **Generate with a strong teacher.** Use the best model you can afford for
generation — quality of the teacher bounds quality of the data. Generate
multiple candidates per seed.
4. **Verify and filter.** Run domain-appropriate checks: execute code, check math
with a symbolic or second-model verifier, use an LLM judge with a strict rubric
for open-ended content. Typical keep rates: 20–60%. Log why examples were
rejected.
5. **Deduplicate and balance.** Near-dup removal (embedding similarity), difficulty
balancing, and format consistency. A synthetic dataset of 10k diverse, verified
examples beats 100k repetitive ones.
6. **Mix with real data.** Blend synthetic with real examples (common starting
ratios: 1:1 to 3:1 synthetic:real). Ablate the ratio — the optimal mix is
task-specific.
7. **Train and measure the delta.** Fine-tune the student, evaluate on held-out
benchmarks for the target capability AND on general benchmarks to catch
regressions. If the delta is zero, fix the pipeline (usually verification or
diversity), not the training hyperparameters.
8. **Record provenance.** Every example tagged with teacher model, prompt version,
filter version. This is what makes the dataset auditable and regenerable.
Checklist for a synthetic dataset release:
- Verification method documented with measured precision on a human-labeled sample.
- Diversity measured (not asserted): embedding coverage, difficulty histogram.
- Real-data mix ratio chosen by ablation, not by default.
- Provenance recorded per example.
- Downstream delta measured on held-out evals, including regression checks.
## Common pitfalls
- **No verification.** The #1 failure. Unverified synthetic data teaches the
student to imitate the teacher's mistakes confidently.
- **Teacher too weak.** A student can't exceed its teacher's verified quality. If
the teacher can't solve the task reliably, its synthetic data is noise.
- **Diversity theater.** Generating 100k examples from 50 seeds gives you 100k
near-duplicates. Seed combinatorics and explicit novelty pressure matter more
than raw count.
- **Format overfitting.** The student learns the teacher's stylistic tics (headers,
"Certainly!") rather than the substance. Vary output formats or strip
boilerplate.
- **Eval contamination.** Synthetic data derived from or resembling benchmark items
inflates scores without improving capability. Keep generation seeds disjoint
from eval content.
- **Ignoring the real-data mix.** Pure synthetic training drifts. Blend with real
data and ablate the ratio — don't guess it.
- **Weak verifier, strong claims.** A sloppy LLM judge passing bad examples
poisons the dataset quietly. Measure your verifier's precision on human-labeled
samples.
- **One-shot pipeline.** Building the pipeline once and never iterating. The first
version's keep rate and diversity are always wrong — instrument, inspect
rejects, improve.