Text Models18 min / updated 2026-09-30

How to Train a Text LLM With LoRA, QLoRA, SFT, DPO, and GRPO

A practical 2026 playbook for adapting open text models with supervised fine-tuning, preference optimization, and reinforcement-style post-training.

Built for: Teams fine-tuning open LLMs for support, coding, data extraction, research assistants, or private domain workflows.

key takeaways

  • +Start with SFT only when you need the model to change behavior, format, or domain language. Use RAG for changing facts.
  • +LoRA and QLoRA are the default first experiment because they keep GPU cost down and preserve the base model.
  • +Preference training matters after SFT when the issue is ranking, refusal style, verbosity, or tool-use behavior.
  • +Use an eval set before training. Otherwise every checkpoint will sound better during manual testing.

workflow - practical training loop

01

Define the production task and freeze a representative eval set.

02

Adapt the model with the smallest method that can change the target behavior.

03

Compare base, checkpoint, and served model before replacing production traffic.

What you are actually training

A text fine-tune does not install a private database into the model. It changes weights so the model is more likely to produce a certain style of answer, follow a format, solve a domain task, or use a tool protocol. For facts that update weekly, retrieval is still the better system design.

In 2026 the practical path is rarely full pretraining. Most teams start with a strong open instruct model, add LoRA adapters, run supervised fine-tuning on high-quality examples, then optionally run DPO or GRPO when there is preference data or verifiable rewards.

  • -Use SFT for: structured output, domain language, code style, customer support tone, tool-call traces, and repeated reasoning patterns.
  • -Use DPO for: choosing the better of two responses when you can label preference pairs.
  • -Use GRPO for: tasks with automatically checkable outcomes, such as math, code tests, exact JSON validity, or simulator rewards.
  • -Use full fine-tuning only when adapter capacity is the bottleneck and you can afford careful regression testing.

Dataset shape that works

The common failure mode is quantity-first data. A 4,000-example dataset where every row is reviewed, deduplicated, and aligned to the production task is usually more valuable than 200,000 scraped conversations.

For SFT, store examples in chat format, keep system prompts explicit, and include negative edge cases. For DPO, create pairs where the rejected answer is plausible enough to teach a boundary. For GRPO, write a reward function that checks the outcome, not whether the answer sounds smart.

  • -Split by source document or customer account, not random row, to prevent leakage.
  • -Hold out tasks that look like production failures: malformed input, hostile prompts, missing context, ambiguous requests.
  • -Track data provenance, license, user consent, PII removal, and whether the data can be used to train a model.
  • -Keep a small golden eval set untouched across all experiments.

Minimal TRL SFT loop

Hugging Face TRL is the most direct path when you want code-level control. PEFT adds LoRA adapters, Transformers loads the base model, and TRL handles the post-training loop.

SFT with TRL and PEFT

python

from datasets import load_dataset
from peft import LoraConfig
from transformers import AutoModelForCausalLM, AutoTokenizer
from trl import SFTConfig, SFTTrainer

model_id = "meta-llama/Llama-3.1-8B-Instruct"
dataset = load_dataset("json", data_files="train.chat.jsonl", split="train")

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    device_map="auto",
    load_in_4bit=True,
)

peft_config = LoraConfig(
    r=32,
    lora_alpha=64,
    lora_dropout=0.05,
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"],
    task_type="CAUSAL_LM",
)

args = SFTConfig(
    output_dir="runs/support-agent-lora",
    per_device_train_batch_size=2,
    gradient_accumulation_steps=8,
    learning_rate=2e-4,
    max_seq_length=4096,
    num_train_epochs=2,
    packing=True,
    eval_strategy="steps",
    eval_steps=100,
)

trainer = SFTTrainer(
    model=model,
    tokenizer=tokenizer,
    args=args,
    train_dataset=dataset,
    peft_config=peft_config,
)
trainer.train()

When to use Axolotl or torchtune instead

Axolotl is excellent when you want reproducible YAML-driven training across many open models. It is especially practical for teams that want to hand configs around, run the same recipe on cloud GPUs, and keep training choices reviewable.

torchtune is better when the team wants PyTorch-native recipes that are readable and hackable. If you are changing model internals, debugging dataloaders, or teaching a team how post-training works under the hood, torchtune is a cleaner mental model than a giant abstraction.

Axolotl-style config sketch

yaml

base_model: meta-llama/Llama-3.1-8B-Instruct
sequence_len: 4096
datasets:
  - path: ./train.chat.jsonl
    type: chat_template
adapter: qlora
lora_r: 32
lora_alpha: 64
lora_dropout: 0.05
gradient_accumulation_steps: 8
micro_batch_size: 2
num_epochs: 2
learning_rate: 0.0002
optimizer: adamw_bnb_8bit
bf16: auto
evals_per_epoch: 4
output_dir: ./outputs/support-agent

Optimization path after the first checkpoint

Do not tune every knob at once. First confirm the model has learned the task. Then search rank, target modules, sequence length, packing, learning rate, and data mixture. A boring experiment log beats heroic guesswork.

If the model overfits style but misses facts, reduce epochs and improve retrieval. If it memorizes refusal language, rebalance the dataset. If it fails JSON or tool calls, add constrained decoding or parser retries instead of expecting weights alone to solve it.

  • -Run evals on base model, SFT checkpoint, merged checkpoint, and served checkpoint.
  • -Keep adapters separate until the eval story is stable.
  • -Measure latency and VRAM after merging because training wins can disappear at serving time.
  • -Release with rollback: base model, previous adapter, and new adapter should all be deployable.

Sources and further reading