Speech Recognition13 min / updated 2026-09-30

Fine-Tuning Whisper and ASR Models for Domain Audio

How to build a labeled speech-recognition dataset, fine-tune Whisper-style ASR, and evaluate word error rate by domain slice.

Built for: Teams transcribing meetings, calls, lectures, podcasts, medical dictation, legal audio, or noisy field recordings.

key takeaways

  • +ASR fine-tuning is most useful for accents, vocabulary, acoustic conditions, or formatting that the base model misses.
  • +Your validation set should represent hard production audio, not just clean studio clips.
  • +Measure WER overall and by slice: speaker, accent, microphone, noise, language, and domain vocabulary.
  • +Keep formatting rules separate from acoustic transcription where possible.

evaluation graph - train what you can measure

data quality5/5
prompt/control fit4/5
temporal consistency4/5
safety and rights5/5

When Whisper fine-tuning helps

OpenAI Whisper was pretrained on a large multilingual supervised corpus, which makes it unusually robust out of the box. Fine-tuning is worth doing when your deployment has repeated errors: specialized vocabulary, low-resource language variants, call-center audio, medical terms, code-switched speech, or custom punctuation requirements.

If the problem is speaker diarization or separating overlapping voices, fine-tuning Whisper alone will not solve it. Add a diarization model and build a pipeline.

Dataset design

Every training row should have an audio file and a transcript that represents what should be emitted. Keep raw transcripts and normalized transcripts if you need both literal speech and clean business output.

  • -Use lossless or high-quality audio during training; avoid repeated MP3 transcodes.
  • -Segment long files into shorter chunks with natural phrase boundaries.
  • -Mark unintelligible spans consistently instead of inventing words.
  • -Create a test set from real production failures and never train on it.

Dataset columns

json

{"audio": "calls/0421_000.wav", "sentence": "The patient reported tachycardia after the second dose."}
{"audio": "calls/0421_001.wav", "sentence": "Schedule a follow up appointment for next Thursday."}

Training skeleton

The Hugging Face workflow uses Datasets for audio loading, Transformers for the Whisper checkpoint, and a sequence-to-sequence trainer. Start with a smaller checkpoint to validate the pipeline before spending money on larger runs.

Whisper fine-tuning outline

python

from datasets import Audio, load_dataset
from transformers import WhisperForConditionalGeneration, WhisperProcessor, Seq2SeqTrainer

model_id = "openai/whisper-small"
dataset = load_dataset("json", data_files={"train": "train.jsonl", "eval": "eval.jsonl"})
dataset = dataset.cast_column("audio", Audio(sampling_rate=16000))

processor = WhisperProcessor.from_pretrained(model_id, language="en", task="transcribe")
model = WhisperForConditionalGeneration.from_pretrained(model_id)

def prepare(batch):
    audio = batch["audio"]
    batch["input_features"] = processor.feature_extractor(
        audio["array"], sampling_rate=audio["sampling_rate"]
    ).input_features[0]
    batch["labels"] = processor.tokenizer(batch["sentence"]).input_ids
    return batch

dataset = dataset.map(prepare, remove_columns=dataset["train"].column_names)

# Add Seq2SeqTrainingArguments, a data collator, WER metric, then train.
trainer = Seq2SeqTrainer(model=model, train_dataset=dataset["train"], eval_dataset=dataset["eval"])
trainer.train()

Evaluation slices

WER is necessary but incomplete. A model can improve average WER while getting worse on the accents, microphones, or compliance vocabulary that matter most. Keep metadata on every validation clip and report sliced metrics.

  • -By acoustic condition: studio, laptop mic, phone, vehicle, crowd, music.
  • -By speaker group: accent, language, gender, age range where ethically collected and consented.
  • -By vocabulary: product names, drug names, legal terms, ticket IDs, addresses.
  • -By output format: punctuation, casing, timestamps, diarization compatibility.

Sources and further reading