LLM Inference Optimization: vLLM, TensorRT-LLM, KV Cache, Quantization, and Speculative Decoding
A deployment guide to making LLM inference faster and cheaper without destroying quality, covering serving engines, batching, cache design, quantization, and decoding.
Built for: Platform teams, AI infrastructure engineers, SaaS teams serving chat, agents, extraction, coding, or high-volume inference workloads.
key takeaways
- +Inference optimization is a queueing problem, not only a model problem.
- +KV cache memory dominates long-context serving and multi-turn chat economics.
- +Continuous batching, prefix caching, quantization, and speculative decoding attack different bottlenecks.
- +Measure time-to-first-token, inter-token latency, throughput, tail latency, cost per successful task, and quality regressions together.
optimization map - match lever to bottleneck
core metrics
TTFT
ITL
p95
$/task
Start with the bottleneck
LLM serving has multiple bottlenecks: prefill compute, decode memory bandwidth, KV cache memory, scheduler efficiency, network overhead, and model quality. The right optimization depends on which one is limiting your workload.
A long-document extraction system and a short customer-support chatbot need different serving choices. Long prompts benefit from prefix caching and prefill optimization. High-concurrency chat benefits from continuous batching and KV cache management.
Serving engine choices
vLLM became popular because PagedAttention and continuous batching make high-throughput serving easier. TensorRT-LLM targets NVIDIA-optimized production inference with graph optimizations, quantization, and advanced decoding. llama.cpp is invaluable for local, CPU, Apple Silicon, and quantized edge deployments.
Optimization map
text
Problem Likely lever
High concurrent chat vLLM continuous batching
Repeated system prompts Prefix caching
Long-context memory pressure KV cache quantization / paging
NVIDIA production serving TensorRT-LLM
Local or edge deployment llama.cpp quantized models
Slow decode Speculative decoding
Expensive simple requests Router to smaller modelKV cache economics
The KV cache stores attention keys and values for tokens already processed. It is what makes generation efficient, but it can consume enormous GPU memory for long context and many concurrent users.
Prefix caching helps when many requests share the same system prompt, policy, tool schema, or retrieved corpus prefix. Cache-aware prompt design can produce real infrastructure savings.
- -Keep common prompt prefixes byte-identical so caches hit.
- -Separate stable system/tool text from per-user context.
- -Track cache hit rate by route, customer, and prompt version.
- -Test long-context tail latency, not only average latency.
Speculative decoding and quantization
Speculative decoding uses a smaller draft model to propose tokens that the larger target model verifies. It can improve decode speed when the draft model is cheap and agrees often enough.
Quantization reduces memory and can improve throughput, but every quantized deployment needs quality regression tests. The failures are often task-specific: JSON accuracy, code correctness, math, rare languages, or refusal boundaries.
- -Use quantization first for high-volume, tolerant tasks like classification or extraction.
- -Be cautious with regulated answers, code generation, and math-heavy workflows.
- -Evaluate served model outputs, not only offline checkpoint quality.
Production metrics
A serving stack is healthy when it meets product latency and reliability goals at an acceptable cost without silent quality loss.
- -TTFT: time to first token.
- -ITL: inter-token latency during decode.
- -Throughput: completed tokens and completed tasks per second.
- -Tail latency: p95 and p99 by route and prompt length.
- -Quality: task pass rate after quantization, routing, and decoding changes.
Sources and further reading
vLLM documentation
Serving engine documentation including PagedAttention, batching, and production features.
TensorRT-LLM documentation
NVIDIA optimized LLM inference framework documentation.
llama.cpp repository
Local and edge inference runtime with quantization and broad hardware support.
FlashAttention-3 paper
Research on faster attention for Hopper GPUs.