Synthetic Data17 min / updated 2026-09-30

Synthetic Data and Model Distillation: How to Generate Training Data Without Poisoning Your Evals

A practical framework for synthetic data generation, teacher-student distillation, data filtering, eval hygiene, and model improvement loops.

Built for: Teams improving smaller models, creating domain datasets, bootstrapping evals, generating tool-use traces, or reducing dependence on proprietary models.

key takeaways

  • +Synthetic data is useful when it expands coverage of a known task, not when it replaces knowing what good data looks like.
  • +Never generate training data from your held-out eval cases or near-duplicates of them.
  • +Use strong teachers, explicit rubrics, filtering, deduplication, and adversarial validation.
  • +Distillation should be judged by task success and cost reduction, not by whether the student imitates every teacher phrase.

data loop - generate, filter, evaluate

01

Teacher model generates candidate examples from audited seed tasks and rubrics.

02

Validators filter schema errors, duplicates, unsafe rows, and eval-set neighbors.

03

Student model trains only on accepted data and is scored against independent evals.

What synthetic data is good for

Synthetic data is most valuable when real examples are scarce but the desired behavior is easy to specify or verify. It can create format diversity, edge cases, tool-call traces, refusal scenarios, multilingual variants, and structured reasoning examples.

It is dangerous when teams use it to paper over unknown requirements. A model trained on synthetic guesses will become confidently aligned to those guesses.

  • -Good: generate SQL question variants for a known schema.
  • -Good: create tool-call traces from audited workflows.
  • -Good: expand rare error cases and edge formats.
  • -Bad: generate medical advice without expert review.
  • -Bad: create eval answers with the same model being evaluated.

Teacher-student workflow

A practical distillation loop uses a stronger teacher model to generate or label examples, filters those examples with rules and model judges, trains a cheaper student model, and evaluates the student on independent human-authored or production-derived tasks.

Distillation loop

text

Seed tasks
  -> teacher generates candidate examples
  -> validators check schema, citations, safety, uniqueness
  -> reviewers inspect samples by slice
  -> student fine-tunes on accepted data
  -> independent eval set compares base, teacher, and student
  -> production canary with rollback

Data filtering

Filtering is where synthetic data projects usually win or lose. Keep the generator and grader separate where possible. Use deterministic checks first, model judgments second, and human review for sensitive slices.

  • -Validate schemas and tool arguments with parsers.
  • -Deduplicate by exact match and embedding similarity.
  • -Reject examples too close to the eval set.
  • -Score diversity by task type, length, language, domain, and failure mode.
  • -Track which teacher model and prompt produced each row.

Synthetic row metadata

json

{
  "task": "extract_invoice_fields",
  "input": "...",
  "output": {"invoice_id": "INV-4017", "total": 1290.50},
  "teacher": "frontier-model-x",
  "generator_prompt_version": "synth-v4",
  "validators": ["json_schema", "currency_check", "dedupe_v2"],
  "accepted": true
}

Eval hygiene

Synthetic data makes eval leakage easier. If the generation prompt has access to eval examples, or if the teacher paraphrases eval cases into the training set, your numbers will look excellent and collapse in production.

  • -Freeze eval sets before large synthetic generation runs.
  • -Compute similarity between synthetic training rows and eval rows.
  • -Keep separate model judges for training-data filtering and final eval grading.
  • -Report performance by synthetic-only, real-only, and mixed eval slices.

Sources and further reading