Inference Optimization18 min / updated 2026-09-30

LLM Inference Optimization: vLLM, TensorRT-LLM, KV Cache, Quantization, and Speculative Decoding

A deployment guide to making LLM inference faster and cheaper without destroying quality, covering serving engines, batching, cache design, quantization, and decoding.

Built for: Platform teams, AI infrastructure engineers, SaaS teams serving chat, agents, extraction, coding, or high-volume inference workloads.

key takeaways

  • +Inference optimization is a queueing problem, not only a model problem.
  • +KV cache memory dominates long-context serving and multi-turn chat economics.
  • +Continuous batching, prefix caching, quantization, and speculative decoding attack different bottlenecks.
  • +Measure time-to-first-token, inter-token latency, throughput, tail latency, cost per successful task, and quality regressions together.

optimization map - match lever to bottleneck

continuous batching5/5
prefix caching4/5
quantization4/5
speculative decoding3/5

core metrics

TTFT

ITL

p95

$/task

Start with the bottleneck

LLM serving has multiple bottlenecks: prefill compute, decode memory bandwidth, KV cache memory, scheduler efficiency, network overhead, and model quality. The right optimization depends on which one is limiting your workload.

A long-document extraction system and a short customer-support chatbot need different serving choices. Long prompts benefit from prefix caching and prefill optimization. High-concurrency chat benefits from continuous batching and KV cache management.

Serving engine choices

vLLM became popular because PagedAttention and continuous batching make high-throughput serving easier. TensorRT-LLM targets NVIDIA-optimized production inference with graph optimizations, quantization, and advanced decoding. llama.cpp is invaluable for local, CPU, Apple Silicon, and quantized edge deployments.

Optimization map

text

Problem                       Likely lever
High concurrent chat           vLLM continuous batching
Repeated system prompts        Prefix caching
Long-context memory pressure   KV cache quantization / paging
NVIDIA production serving      TensorRT-LLM
Local or edge deployment       llama.cpp quantized models
Slow decode                    Speculative decoding
Expensive simple requests      Router to smaller model

KV cache economics

The KV cache stores attention keys and values for tokens already processed. It is what makes generation efficient, but it can consume enormous GPU memory for long context and many concurrent users.

Prefix caching helps when many requests share the same system prompt, policy, tool schema, or retrieved corpus prefix. Cache-aware prompt design can produce real infrastructure savings.

  • -Keep common prompt prefixes byte-identical so caches hit.
  • -Separate stable system/tool text from per-user context.
  • -Track cache hit rate by route, customer, and prompt version.
  • -Test long-context tail latency, not only average latency.

Speculative decoding and quantization

Speculative decoding uses a smaller draft model to propose tokens that the larger target model verifies. It can improve decode speed when the draft model is cheap and agrees often enough.

Quantization reduces memory and can improve throughput, but every quantized deployment needs quality regression tests. The failures are often task-specific: JSON accuracy, code correctness, math, rare languages, or refusal boundaries.

  • -Use quantization first for high-volume, tolerant tasks like classification or extraction.
  • -Be cautious with regulated answers, code generation, and math-heavy workflows.
  • -Evaluate served model outputs, not only offline checkpoint quality.

Production metrics

A serving stack is healthy when it meets product latency and reliability goals at an acceptable cost without silent quality loss.

  • -TTFT: time to first token.
  • -ITL: inter-token latency during decode.
  • -Throughput: completed tokens and completed tasks per second.
  • -Tail latency: p95 and p99 by route and prompt length.
  • -Quality: task pass rate after quantization, routing, and decoding changes.

Sources and further reading