Training Music Generation Models With AudioCraft, MusicGen, and Stable Audio
How to prepare licensed music datasets, fine-tune text-to-music models, and evaluate style, structure, BPM, and legal risk.
Built for: Music-tech teams, game studios, sample-pack creators, and researchers building controllable music generation.
key takeaways
- +Music data rights matter more than almost any modeling choice.
- +Captioning should include genre, instrumentation, BPM, key, arrangement, mood, and production texture.
- +Fine-tuning can teach a palette or production vocabulary, but structural composition remains hard.
- +Evaluate musical usefulness, not just audio realism.
evaluation graph - train what you can measure
Model families
MusicGen models music as discrete audio tokens with text or melody conditioning, using EnCodec-style tokenization under the AudioCraft stack. Stable Audio approaches text-to-audio with diffusion-style tooling and has open tooling for inference and fine-tuning workflows.
For production, the choice is often less about benchmark quality and more about rights, controllability, latency, output duration, and whether commercial use is allowed.
Dataset preparation
Do not train on random streaming rips. Build or license a dataset where training rights are explicit. For every clip, store source, license, contributor, BPM, key, instrumentation, genre, mood, and whether vocals are present.
Clip length should match the product. A one-shot sample generator needs different data than a 90-second background music generator.
- -Normalize loudness but keep dynamics; over-compression teaches flat outputs.
- -Separate loops, one-shots, stems, and full mixes.
- -Caption arrangement: intro, drop, build, fill, ending, sparse verse, dense chorus.
- -Hold out artists, sample packs, or albums by source to avoid leakage.
Music caption row
json
{"audio": "licensed/loop_081.wav", "caption": "124 BPM deep house drum loop, warm kick, tight closed hats, no vocals, four bar seamless loop", "bpm": 124, "key": null, "license": "internal-training-ok"}
{"audio": "licensed/cue_014.wav", "caption": "cinematic ambient cue in D minor, bowed strings, soft piano, slow rise, no percussion", "bpm": 72, "key": "D minor", "license": "internal-training-ok"}Fine-tuning workflows
AudioCraft uses Dora configs for MusicGen training and fine-tuning. Stable Audio tooling uses checkpoint continuation and config files. In both cases, run a tiny overfit test first: the model should be able to learn a handful of clips before you launch an expensive training run.
AudioCraft and Stable Audio command shapes
bash
# MusicGen fine-tune shape
dora run solver=musicgen/musicgen_base_32khz model/lm/model_scale=medium continue_from=//pretrained/facebook/musicgen-medium conditioner=text2music
# Stable Audio continuation shape
python train.py --config-file ./defaults.ini --pretrained-ckpt-path ./checkpoints/stable-audio-open.ckpt --save-dir ./runs/brand-palette --batch-size 4 --precision 16Evaluation for music usefulness
A model that produces plausible audio can still be useless for musicians. Evaluate whether outputs loop cleanly, preserve BPM, avoid unwanted vocals, obey instrumentation prompts, and can be edited in a DAW.
- -BPM adherence: compare detected tempo against requested tempo.
- -Loopability: check transient continuity at loop boundaries.
- -Prompt control: genre, instrument, key, intensity, vocals/no vocals.
- -Novelty and leakage: similarity search against training clips.
- -Production value: noise floor, clipping, stereo image, mix balance.
Sources and further reading
AudioCraft repository
MusicGen and AudioGen inference and training code.
AudioCraft MusicGen docs
MusicGen pretraining and fine-tuning command patterns.
MusicGen paper
Simple and Controllable Music Generation research paper.
Stable Audio tools
Fine-tuning flags and checkpoint continuation workflow.
Stable Audio 3 repository
Current open platform for audio/music inference and LoRA fine-tuning.