Llm Genai Ops

LLM Inference Optimization: Serving More Tokens Per GPU

0 of 11 complete

0%

Contents

Back|Llm Genai OpsLLM Inference Optimization: Serving More Tokens Per GPU
1/11
55 min left
  1. Home
  2. AI Engineering: Evaluation, LLM Ops and Security
  3. LLM and GenAI Ops
  4. LLM Inference Optimization: Serving More Tokens Per GPU
Prerequisites
Vector Databases and Approximate Nearest Neighbor Searchrequired
Related Topics
Scaling and GPU Infrastructure: Serving Models Without Burning MoneyCore ConceptsWhy ML Models Fail in Production: The Production GapCore ConceptsModel Packaging and Containerization: Killing 'Works On My Machine' for MLCore ConceptsModel Serving and Inference APIs: Turning a Model File Into a ServiceCore ConceptsTraining Pipelines and Orchestration: Retraining That Runs ItselfCore Concepts
Previous lessonVector Databases and Approximate Nearest Neighbor SearchNext lesson
1 of 11
LLM Guardrails and Safety: Wrapping the Model So It Can Ship

System Design

  • Foundation
  • Intermediate
  • Advanced
  • Capstone

AI Engineering

  • Foundation
  • Data, RAG and Agents
  • Evaluation, LLM Ops and Security

systemdesign.academy

  • Home
  • Glossary
  • Interview prep
  • Reviews
  • About
  • Privacy
  • Terms

Where this shows up in interviews

4 system design questions lean on this idea. Each walks through the full answer.

  • →Design a Recommendation System
  • →Design an LLM Inference Platform
  • →Design ChatGPT
  • →Design RAG System

The Bill That Ends Most GenAI Projects

A team ships a chatbot. The demo is great. Then the first real month of traffic arrives, and finance forwards the GPU invoice. A single NVIDIA H100 rents for roughly two to four dollars an hour. Their naive server, running one request at a time, handles maybe a handful of users per card. To serve real traffic they now need a rack of these things, and the math turns a promising feature into a project nobody can afford.

This is the quiet killer of GenAI in production. The model works. The problem is that serving it is slow and expensive in a way that ordinary web services are not. A REST endpoint returns in a few milliseconds and one server handles thousands of users. An LLM chews on each request for seconds and one expensive GPU handles a few users at a time, unless you optimize it hard.

Here is the good news, and it is the whole reason this lesson exists. The gap between a naive server and a tuned one is not 20 percent. It is often 10 to 20 times more throughput on the exact same hardware, which means 10 to 20 times lower cost per token with no change to the model at all. That gap is entirely engineering: the , continuous batching, PagedAttention, quantization, and speculative decoding. Every one of them is a lever you control, and by the end of this lesson you will know what each does, roughly what it buys, and what it costs.

The model is fixed. Your GPU bill is not. It is set by how well you serve, and that is a platform problem, not a modeling one.

We will work from the ground up. First the one hardware fact that explains everything, then the memory that dominates the card, then the batching and memory tricks that deliver most of the win, then and speculative decoding on top, and finally the runtimes that package all of it so you do not write a line of it yourself.

Why LLM Inference Is Fundamentally Slow

To fix inference you have to understand why it is slow in the first place, and the answer is the word autoregressive. A large language model generates text one token at a time, and each new token depends on all the tokens before it. It cannot jump ahead. To write a 500-token answer it must run the full model 500 times in sequence. There is no way to parallelize across the tokens of a single reply, because token 200 literally does not exist until token 199 has been produced and fed back in.

Now the surprising part, and the fact the rest of the lesson hangs on. Each of those steps is not compute-bound, it is memory-bound. On every step the GPU has to read the entire set of model weights out of VRAM to produce a single token. For a 13-billion-parameter model in FP16 that is 26 gigabytes of reading, per token. The GPU's thousands of math units sit mostly idle, starved, waiting on the memory system. The bottleneck is memory , how fast you can move weights out of VRAM, not raw floating-point power.

Engineers formalize this with a roofline. A device only reaches its peak math rate when a workload does enough arithmetic per byte it reads. Decode does almost none, so it lives on the sloped, bandwidth-limited side. Prefill, which processes the whole prompt at once, reuses each weight across hundreds of tokens and climbs into the flat, compute-limited region.

Roofline chart contrasting decode and prefill on an H100-class GPU: the x-axis is arithmetic intensity in FLOP per byte read on a log scale and the y-axis is achievable throughput as a percent of peak, with a shaded memory-bound region on the left where decode sits at roughly 0.5 FLOP per byte far below the ceiling, and a compute-bound region on the right where prefill sits at roughly 400 FLOP per byte on the flat peak, showing decode is starved by bandwidth while prefill is not

That single fact explains almost every optimization in this lesson:

  • If decode is limited by moving bytes, then shrinking the weights () makes it faster, because there are fewer bytes to move.
  • If you are reading the weights anyway, then applying them to many requests at once (batching) is nearly free extra , because the expensive read is already paid for.
  • If each slow step only makes one token, then making the big model produce several tokens per step (speculative decoding) is a direct win against the sequential bottleneck.

Hold onto the roofline as you read on. Every technique below is either moving fewer bytes or wringing more useful work out of each byte you were already forced to read.

Prefill vs Decode, TTFT vs Tokens Per Second

Before the optimizations, split a single request into its two phases, because they behave nothing alike and confusing them will mislead every performance decision you make.

Two-phase request timeline: a wide blue PREFILL block on the left labelled 512 tokens in one parallel pass and compute-bound, a dashed marker at 180 milliseconds where the first token appears marking time to first token, then a long run of small green DECODE ticks each producing one token per 20 millisecond step labelled memory-bound and sequential, over a wall-clock time axis, with two metric cards below reporting TTFT of 180 milliseconds and tokens per second of 50

Prefill processes the entire prompt at once. All 512 prompt tokens go through the model in one big parallel pass. This phase is compute-bound, it lives in the healthy part of the roofline, and it is what you wait on before the first word appears. A longer prompt means a slower start, because there is more prompt to push through that first pass.

Decode then generates the answer one token at a time, feeding each output back in as the next input. This phase is the memory-bound, sequential grind, and it sets how fast the text streams out after it begins.

Those two phases map onto the two numbers you actually report to users and to your capacity planning:

MetricWhat it measuresSet mostly byWhat the user feels
Time to first token (TTFT)Delay before the first word appearsPrompt length and prefill compute"How long until it starts?"

The KV Cache: The Memory That Runs Everything

If a model must look at every previous token to produce the next one, then naively it would recompute the attention over the entire history on every single step. Step 100 would redo the work of steps 1 through 99, step 101 would redo 1 through 100, and so on. That would make generation quadratically slow in the length of the output, which no production system could tolerate.

Instead, every model uses a . Attention works by turning each token into a key and a value. Once you have computed the key and value for a token, they never change, because the token behind them never changes. So you compute them once and cache them in VRAM. On each new step the model only computes the key and value for the one new token and reuses everything cached for the past. The quadratic recomputation problem collapses to linear work.

But this cache has a cost, and it is the single most important number in LLM serving: the KV cache lives in VRAM, and it grows with every token, for every concurrent request. For a model like Llama-2 13B, the KV cache runs on the order of 1 megabyte per token of context. Multiply that by a few thousand tokens of context and by dozens of concurrent users and the cache can easily eat more VRAM than the model weights themselves.

KV cache growth line chart on an 80 GB H100 running a 13B model in INT8: the x-axis is concurrent users each at 4096 tokens of context and the y-axis is VRAM used in gigabytes, a purple KV cache line climbs linearly and crosses the flat 13 GB weights line at only four users where 16 GB of KV cache already exceeds the weights, heading toward the 80 GB card ceiling, showing the KV cache, not the weights, is what fills the card

Read the crossover on that chart. The weights are a flat, one-time 13 GB. The KV cache is a rising line that passes the weights at just four concurrent users and keeps climbing toward the ceiling. Long before compute becomes the limit, you run out of room to hold everyone's cache.

Newer models attack this at the architecture level with grouped-query attention, where several attention heads share one set of keys and values instead of each keeping its own. Llama-3 and Mistral use it, and it shrinks the cache severalfold. Even then the KV cache stays the dominant pressure on VRAM, which is why VRAM, not compute, is usually what limits how many users one GPU can serve. How you manage that cache decides your , and it is exactly what the next two techniques are fighting over.

Continuous Batching and PagedAttention: The 10x

Since decode reads the weights from VRAM on every step regardless of how many requests are in flight, you may as well apply those weights to many sequences in the same pass. That is batching, and it is where most of the 10 to 20x lives. But for LLMs there is a right way and a wrong way to do it, and the difference is the whole ballgame.

The wrong way is static batching: gather a group of requests, run them together, and hold the whole group until the last one finishes. The problem is that reply lengths vary by 100x. A one-line answer gets stuck waiting for a 2000-token essay in the same batch, and its slot on the GPU sits idle the entire time it waits. Utilization craters to a third of the card or worse, and you are paying full price for a mostly empty GPU.

The right way is , popularized by vLLM. Rebuild the batch on every single decode step. The moment a sequence finishes it leaves and its slot is freed. New requests join on the very next step instead of waiting for a batch boundary. The GPU is always working on a full batch of active sequences, so the reads you were already paying for get spread across as many users as the card can hold.

Static versus continuous batching Gantt timeline across four GPU slots over 20 decode steps: in the static half a short reply R1 finishes at step 4 but its slot stays hatched red and idle until the longest reply R4 ends at step 20, giving about 34 percent GPU utilization, while in the continuous half a freed slot immediately admits new requests R5 and R6 so the slots stay full, giving about 92 percent GPU utilization, illustrating that the reclaimed idle time is the source of the throughput gain

The hatched red in that figure is paid-for GPU doing nothing. Static batching leaves the short replies holding dead slots; continuous batching reclaims each slot the instant a reply ends and slides a waiting request in. That reclaimed idle time, not any change to the model, is where the multiple comes from.

Continuous batching only works if you can hand out and reclaim slots cheaply, and that is what PagedAttention adds. Older servers stored each sequence's KV cache as one contiguous block sized for the maximum possible length. A short reply left most of its reservation empty, and that empty space could not be lent to anyone else. The vLLM paper (Kwon et al., SOSP 2023) measured how much this cost: "only 20.4% - 38.2% of the KV cache memory is used to store the actual token states in the existing systems." The rest was reserved for tokens not yet generated, or lost to internal and external fragmentation.

Quantization: Fewer Bits, Smaller Bill

Decode is memory-bound, so the most direct way to speed it up is to make the weights smaller, because fewer bytes to read is less time reading. stores each weight in fewer bits. FP16 is the 16-bit baseline, INT8 halves it, and INT4 quarters it. Fewer bits means less VRAM to hold the model and fewer bytes to move every step, which is a direct decode speedup, and it also frees room for the that actually limits concurrency. Cutting a 13B model from 26 GB to 6.5 GB hands almost 20 GB back to the cache on an 80 GB card.

Quantization precision comparison for a 13B model across three cards, FP16, INT8, and INT4: FP16 uses 26 GB of model VRAM leaving 54 GB of KV headroom on an 80 GB card at reference quality, INT8 halves the model to 13 GB leaving 67 GB of headroom at a quality drop under one percent, and INT4 with GPTQ or AWQ quarters it to 6.5 GB leaving 73.5 GB of headroom at a quality drop of about three percent on hard reasoning tasks, showing memory shrinks fast while quality falls slowly

The two names you will meet most for 4-bit are GPTQ and AWQ. Both are calibration-based: they run a small sample of data through the model and choose the rounding so that the weights that matter most to the output keep their precision. That is how you drop to 4 bits and still keep most of the quality. Naive rounding to 4 bits, treating every weight as equally important, would wreck the model, and that difference is why calibration methods exist.

Serving a quantized model in vLLM is a one-line change:

# Serve a 13B model quantized to INT4 with AWQ.
# Fits in ~6.5 GB of VRAM instead of ~26 GB, leaving far more
# room for the KV cache (which is what limits concurrency).
vllm serve TheBloke/Llama-2-13B-chat-AWQ \
  --quantization awq \
  --max-model-len 4096 \
  --gpu-memory-utilization 0.90

# The endpoint is OpenAI-compatible, so clients do not change:
#   POST /v1/chat/completions  { "model": "...", "stream": true }

The trade-off is quality, and it is not linear, which is the useful part. INT8 is nearly free: usually well under a one percent drop, safe to default to for almost any workload. INT4 costs a few percent on the hardest reasoning, math, and code tasks, while staying fine for chat and summarization. The rule is simple. Quantize as aggressively as your own evaluation set will tolerate, and never trust the memory savings without measuring the accuracy on your actual task, because a benchmark number from someone else's task tells you nothing about yours.

Speculative Decoding: Two Models Racing

and quantization make each step cheaper. Speculative decoding attacks a different weakness entirely: one slow model pass only buys you one token. What if a single pass could produce several?

The trick uses two models. A small, fast draft model guesses the next few tokens ahead. Then the big, accurate target model verifies all of those guesses in a single parallel forward pass. Verifying 4 tokens costs almost the same as generating 1, because that pass is memory-bound and you are reading the big weights exactly once either way. Every guess the target agrees with is a token you got nearly for free.

Speculative decoding as a UML sequence diagram with two lifelines, a small fast draft model and a big accurate target model: the draft proposes four tokens the cat sat quietly, the target verifies all four in one parallel pass at roughly the cost of one normal decode step, accepts the longest agreed prefix the cat sat, and at the first mismatch replaces quietly with its own token on, and a result strip below shows three accepted tokens, one rejected, and one target correction, advancing four tokens in a single target pass with output identical to the target alone

The beautiful part is correctness. At the first token where the draft and target disagree, you throw away the rest of the guesses and use the target model's own token, which it already computed during that verification pass. The output is bit-for-bit identical to running the big model alone. There is no quality trade-off whatsoever, which is rare among optimizations. On predictable text, where the little model is often right, this delivers 2 to 3 times faster generation. The only cost is running a second small model and a bit more VRAM to hold it.

Here is a runnable toy that simulates the acceptance loop. There is no real model in it, just a known target sequence and a slightly wrong draft, so you can watch how many expensive target passes it takes and confirm the output never differs from the target.

The output shows fewer target passes than tokens produced, and it confirms the generated text matches the target exactly. The toy accepts a long run and overstates the real gain, which on production text lands at 2 to 3 times, but the mechanism is the point: the expensive model runs fewer times, and it never once compromises the answer.

The Batch-Size Dial: Throughput Against Latency

With the mechanisms in hand, the tuning question becomes concrete. lets you run many sequences at once, but how many should you run? Batch size is the one knob that trades a single user's speed for the whole fleet's efficiency, and there is no universally right setting.

Throughput versus latency dual-axis line chart for a 13B model on one H100: the x-axis is batch size from 1 to 64 concurrent sequences, a blue aggregate tokens per second line rises steeply from about 42 to 1250 tokens per second on the left axis while a red per-user tokens per second line falls from about 42 to 20 on the right axis, showing that growing the batch multiplies total throughput but slows each individual user

Read the two curves. As batch size grows, aggregate climbs steeply, because you reuse each expensive weight read across more sequences. But per-user streaming speed drifts down, because more sequences now share each pass and each one advances a little slower. The right batch size is therefore a product decision, not a technical maximum. An interactive chat caps the batch to protect per-user TPS so the text still types out briskly. An offline job pushes the batch to the limit and does not care that any single item is slow, because it only wants the cheapest possible bulk tokens. The clean way to set it is to pick a per-user TPS floor as your service level, then raise the batch until you are about to breach it.

That efficiency is not academic, it is the entire economic case. Because the GPU rents for a fixed hourly price, cost per token is set entirely by how many tokens per second you extract from the card. Stack the levers from this lesson and the cost falls by the same factor throughput rises.

Cost per million tokens bar chart on one H100 at three dollars an hour for a 13B model, stacking optimizations from top to bottom: a naive server at 40 tokens per second costs 20.83 dollars per million tokens, adding continuous batching with PagedAttention at 560 tokens per second drops it to 1.49 dollars, adding INT4 quantization at 900 tokens per second drops it to 0.93 dollars, and adding speculative decoding at 1600 tokens per second drops it to 0.52 dollars, roughly a fortyfold reduction with no change to the model

The same H100 goes from about 20 dollars per million tokens on a naive server to well under a dollar once the stack is in place. Continuous batching does the heavy lifting, frees VRAM for still more concurrency, and speculative decoding stacks on top for predictable text. Nothing about the model changed. That spread is the difference between a feature nobody can afford and a service that pays for itself, and it is why serving, not modeling, is where the money is made or lost.

The Serving Runtimes, and How Teams Choose

You almost never write this machinery yourself. You pick a serving runtime that has it built in, put it behind an orchestrator, and spend your time tuning rather than implementing. Three runtimes matter, and they overlap heavily, so the real decision is ecosystem fit and how much build complexity you will trade for the last few percent of performance.

Runtime comparison matrix of vLLM, TGI, and TensorRT-LLM across six properties, origin, best for, quantization support, API surface, setup cost, and when to reach for it: vLLM is the fastest to stand up with an OpenAI-compatible API and low setup cost, TGI is the Hugging Face server that fits stacks already on the Hub, and TensorRT-LLM is NVIDIA's compiled runtime for one model at massive scale with a heavier build step

vLLM is the open-source default and the project that made and PagedAttention mainstream. It is the fastest thing to stand up, speaks the OpenAI API out of the box so your clients do not change, and supports GPTQ and AWQ quantization. Most teams start here and never need anything else.

TGI (Text Generation Inference) is Hugging Face's server, tightly integrated with the Hub and a natural fit if your whole stack already lives in the Hugging Face ecosystem. It offers similar continuous batching and support, so the choice between it and vLLM is mostly about which ecosystem you are already invested in.

TensorRT-LLM is NVIDIA's compiled runtime. It squeezes out the last drop of performance by compiling the model into optimized CUDA kernels for a specific GPU, which can beat the others on raw and throughput. The cost is a heavier build step and less flexibility, so teams reach for it when they are serving one model at massive scale and every millisecond counts.

Now the whole stack in one picture. Follow the numbered path from a client request to a streamed token and notice where each optimization from this lesson lives: the scheduler is where continuous batching runs, the KV block manager is where PagedAttention runs, and the model executor is where prefill and decode happen on the quantized weights.

vLLM-style serving architecture on Kubernetes with a numbered request path: a client calls an OpenAI-compatible endpoint, the vLLM server contains an API server, a scheduler running continuous batching, a KV block manager running PagedAttention, and a model executor, all driving an H100 GPU that holds the quantized weights and the paged KV cache, with callouts explaining prefill sets TTFT, the decode loop packs all active sequences each step, and finished sequences return their KV pages so a waiting request takes the slot

Every Lever on One Card

Here is the whole toolkit at a glance: what each lever attacks, roughly what it buys, and the honest cost. Read the last column, because none of these is free, and knowing the cost is what keeps you from cargo-culting an optimization your workload does not need.

Optimization levers taxonomy table with four columns, lever, what it targets, typical gain, and honest cost, across six rows: the KV cache targets decode and turns quadratic work linear at the cost of VRAM that grows per token per user, continuous batching targets decode for a 10 to 20x throughput gain but needs cheap KV slot reclaim, PagedAttention targets memory for several times more users at the cost of a custom kernel built into the runtime, quantization targets decode for a 2 to 4x smaller model at a few percent quality on hard tasks, speculative decoding targets latency for 2 to 3x with a second draft model, and grouped-query attention targets the KV footprint at model design time

Notice the shape of the toolkit. The is not optional, every server runs it, and it is the reason concurrency is a memory question at all. Continuous batching and PagedAttention are the two that deliver the headline 10 to 20x, and they are a package because batching needs the cheap slot reclaim that paging provides. Quantization stacks a memory and speed win on top and buys back VRAM for even more concurrency. Speculative decoding is the pure latency play with no quality cost. Grouped-query attention is chosen when the model is designed, not when you serve it, but it changes how much cache pressure you inherit.

These levers stack, and a production feature usually runs several at once: a quantized model under with PagedAttention on vLLM, with max-model-len and gpu-memory-utilization tuned so the KV cache never runs dry at peak. Reach for each one deliberately. Continuous batching and paging you want almost always. you push as far as your eval set tolerates. Speculative decoding you add when latency matters and your text is predictable enough for the draft to be right often. That deliberate stack, not any single trick, is what closes the gap between a demo and an affordable product.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

Why is the decode phase of LLM inference described as memory-bound rather than compute-bound?

Q2

What is the main advantage of continuous batching over static batching for LLM serving?

Q3

Why does speculative decoding speed up generation without changing the output?

Q4

On an 80 GB GPU serving a 13B model, why does the KV cache, not the weights, usually decide how many users you can serve?

Tokens per second (TPS)
Streaming speed after it starts
Memory and batch size in decode
"How fast does it type?"

You tune these separately, and the levers pull in different directions. A long system prompt hurts TTFT but not TPS. A bigger batch improves total and can hurt per-user TPS a little. A chat UI needs a snappy TTFT above all, because a user is watching a blank box. A batch summarization job does not care about TTFT at all and only wants maximum total tokens per second across the whole fleet. Know which one your product is optimizing for before you touch a single knob, because a server tuned for one is the wrong server for the other.

Two things fight over the same VRAM: the model weights (fixed) and the KV cache (grows with every user and every token). The winner of that fight sets your capacity.

PagedAttention memory-layout comparison: the top row shows contiguous reservation where three sequences each reserve a 12-cell block but use only 3, 5, and 2 cells with the rest hatched as wasted reserved memory, tagged only 20.4 to 38.2 percent of KV cache memory holding token states, as measured in the vLLM paper, and the bottom row shows paged blocks where the KV cache is packed into small fixed-size pages handed out on demand and returned when a sequence finishes, tagged 96.3 percent of KV cache memory holding token states in vLLM, from the same paper

PagedAttention borrows the operating system's oldest trick: it stores the KV cache in small fixed-size blocks, like virtual memory pages, and hands them out on demand. A sequence gets only the pages it currently needs and returns them the moment it finishes. No fragmentation, no over-reservation, so you can pack far more concurrent sequences into the same card. Continuous batching keeps the compute full and paging keeps the memory full, and together they are the reason vLLM routinely delivers many times the throughput of a naive server on identical hardware.

How real teams put it together: a company like DoorDash or Uber serving an internal LLM feature will typically run vLLM on a pool of GPUs behind a Kubernetes service, quantize to INT8 or INT4 to fit more model and more per card, lean on continuous batching to keep the GPUs near full utilization, and tune max-model-len and gpu-memory-utilization so the KV cache never runs out under peak load. That combination is what turns an unaffordable rack of idle GPUs into a service that pays for itself. The model was never the hard part. Serving it economically is.