Skip to content
Back to skills

Synthetic Data

ASecurity

Generate high-quality training data with models — pipelines, filtering, and validation for synthetic datasets that actually help.

  • 2 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 29, 2026
ai-agentsgoperformance

Security analysis

A100/100

Scanned September 29, 2026

npx -y skills add aicodedecode/awesome-muse-skills --skill synthetic-data --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Synthetic Data?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Synthetic Data
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-synthetic-data/badge)](https://www.skillsdirectory.com/skills/aicodedecode-synthetic-data)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: synthetic-data
description: Generate high-quality training data with models — pipelines, filtering, and validation for synthetic datasets that actually help.
category: ai-research
---

## Overview

Synthetic data is training data generated by models rather than collected from
humans or the world: instruction-response pairs, reasoning traces, code with
tests, preference judgments, tool-use trajectories. It's the engine behind most
modern post-training — distillation, self-improvement loops, and domain adaptation
all run on synthetic data. Done well, it converts a strong teacher model's
capabilities into targeted training signal for a student. Done poorly, it produces
fluent garbage that teaches the student the teacher's failure modes.

The core discipline is treating synthetic generation as a data pipeline with
quality gates, not as "prompt a big model and hope." Every synthetic dataset needs:
a generation step with controlled diversity, a verification step that filters for
correctness (execution, entailment checks, cross-model agreement), and a validation
step measuring downstream effect on the student. The failure mode to fear is model
collapse — training on unfiltered synthetic data degrades quality over
generations — which is why verification and mixing with real data matter.

Synthetic data is leverage: it multiplies a strong teacher into arbitrary
quantities of targeted training signal. Like all leverage, it amplifies errors
too. The pipeline's job is to keep the signal while discarding the errors.

## When to use

- Post-training a model for a domain where human data is scarce or expensive
  (specialized reasoning, low-resource languages, niche tools).
- Distilling capabilities from a stronger teacher into a smaller, cheaper student.
- Generating instruction-tuning data at a scale humans can't match.
- Creating adversarial or edge-case examples to patch specific weaknesses found in
  evals.
- Bootstrapping preference data (pairs judged by a strong model) for DPO or
  reward modeling.
- Augmenting thin slices of real data (rare classes, edge cases) where collection
  is impractical.

## Core concepts

- **Generation with diversity control.** Naive sampling collapses to the teacher's
  modal outputs. Use temperature variation, persona/topic seeding, and explicit
  diversity prompts ("generate a problem unlike these examples") to cover the
  space.
- **Verification beats volume.** The highest-leverage step: filter generated data
  for correctness. Executable domains (code, math) verify by running; open-ended
  domains use entailment checks, multi-sample consistency, or a second model as
  judge.
- **Rejection sampling / best-of-N.** Generate N candidates, keep the ones that
  pass verification. This converts compute into quality and is the workhorse of
  synthetic pipelines.
- **Distillation vs. self-improvement.** Distillation transfers from a stronger
  teacher; self-improvement (STaR-style) has the model generate its own rationales
  and keeps the ones leading to correct answers. Both need the same verification
  discipline.
- **Model collapse.** Repeated training on unfiltered model outputs narrows the
  distribution and amplifies errors. Mitigations: keep a substantial fraction of
  real data, filter aggressively, and regenerate from the strongest available
  teacher each round.
- **Contamination hygiene.** Synthetic eval-adjacent data must never leak into
  benchmarks. Track provenance: which teacher, which prompt, which filter version
  produced every example.
- **Teacher-student gap.** The student learns the teacher's verified outputs, not
  the teacher's capabilities. A student trained on traces can match the teacher's
  task performance without matching its generality — know which one you need.
- **Format diversity.** Varying output formats during generation prevents the
  student from overfitting to one teacher's stylistic template.

## Practical workflow

1. **Define the target capability precisely.** "Better reasoning" is not a spec.
   Write the task distribution: input types, difficulty mix, output format. The
   generation prompts derive from this spec.
2. **Seed for diversity.** Build a seed pool of topics, personas, difficulty
   levels, and edge cases. Sample seeds combinatorially so generation covers the
   space rather than clustering.
3. **Generate with a strong teacher.** Use the best model you can afford for
   generation — quality of the teacher bounds quality of the data. Generate
   multiple candidates per seed.
4. **Verify and filter.** Run domain-appropriate checks: execute code, check math
   with a symbolic or second-model verifier, use an LLM judge with a strict rubric
   for open-ended content. Typical keep rates: 20–60%. Log why examples were
   rejected.
5. **Deduplicate and balance.** Near-dup removal (embedding similarity), difficulty
   balancing, and format consistency. A synthetic dataset of 10k diverse, verified
   examples beats 100k repetitive ones.
6. **Mix with real data.** Blend synthetic with real examples (common starting
   ratios: 1:1 to 3:1 synthetic:real). Ablate the ratio — the optimal mix is
   task-specific.
7. **Train and measure the delta.** Fine-tune the student, evaluate on held-out
   benchmarks for the target capability AND on general benchmarks to catch
   regressions. If the delta is zero, fix the pipeline (usually verification or
   diversity), not the training hyperparameters.
8. **Record provenance.** Every example tagged with teacher model, prompt version,
   filter version. This is what makes the dataset auditable and regenerable.

Checklist for a synthetic dataset release:
- Verification method documented with measured precision on a human-labeled sample.
- Diversity measured (not asserted): embedding coverage, difficulty histogram.
- Real-data mix ratio chosen by ablation, not by default.
- Provenance recorded per example.
- Downstream delta measured on held-out evals, including regression checks.

## Common pitfalls

- **No verification.** The #1 failure. Unverified synthetic data teaches the
  student to imitate the teacher's mistakes confidently.
- **Teacher too weak.** A student can't exceed its teacher's verified quality. If
  the teacher can't solve the task reliably, its synthetic data is noise.
- **Diversity theater.** Generating 100k examples from 50 seeds gives you 100k
  near-duplicates. Seed combinatorics and explicit novelty pressure matter more
  than raw count.
- **Format overfitting.** The student learns the teacher's stylistic tics (headers,
  "Certainly!") rather than the substance. Vary output formats or strip
  boilerplate.
- **Eval contamination.** Synthetic data derived from or resembling benchmark items
  inflates scores without improving capability. Keep generation seeds disjoint
  from eval content.
- **Ignoring the real-data mix.** Pure synthetic training drifts. Blend with real
  data and ablate the ratio — don't guess it.
- **Weak verifier, strong claims.** A sloppy LLM judge passing bad examples
  poisons the dataset quietly. Measure your verifier's precision on human-labeled
  samples.
- **One-shot pipeline.** Building the pipeline once and never iterating. The first
  version's keep rate and diversity are always wrong — instrument, inspect
  rejects, improve.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…