Core

Scaling and GPU Infrastructure: Serving Models Without Burning Money

0 of 12 complete

0%

Contents

Back|CoreScaling and GPU Infrastructure: Serving Models Without Burning Money
1/12
55 min left
  1. Home
  2. AI Engineering: Foundation
  3. Core Concepts
  4. Scaling and GPU Infrastructure: Serving Models Without Burning Money
Prerequisites
Monitoring and Drift Detection: Catching a Model That Fails Without an Errorrequired
Related Topics
LLM Inference Optimization: Serving More Tokens Per GPULLM and GenAI OpsFine-Tuning vs RAG vs Prompting: Choosing Your ApproachLLM and GenAI OpsParameter-Efficient Fine-Tuning: LoRA and QLoRALLM and GenAI OpsEvaluating LLMs in Production: Grading Answers That Have No Right AnswerLLM and GenAI OpsPrompt Management and Versioning: Treat Prompts as Production CodeLLM and GenAI Ops
Previous lessonMonitoring and Drift Detection: Catching a Model That Fails Without an Error
1 of 12
Next lesson
Production RAG and LLM Systems: Running It, Not Just Building It

System Design

  • Foundation
  • Intermediate
  • Advanced
  • Capstone

AI Engineering

  • Foundation
  • Data, RAG and Agents
  • Evaluation, LLM Ops and Security

systemdesign.academy

  • Home
  • Glossary
  • Interview prep
  • Reviews
  • About
  • Privacy
  • Terms

Where this shows up in interviews

This idea carries a full system design question on its own. Each walks through the full answer.

  • →Design an LLM Inference Platform

The $40,000 Bill for a Model Nobody Was Using at 3 AM

A team ships a language model behind a chat feature. It works. is fine. Everyone moves on. A month later finance forwards the cloud bill with a note that just says "what is this." The line item is a fleet of GPU instances running twenty-four hours a day, seven days a week, at a few dollars per GPU per hour. Do that math on eight cards and you clear forty thousand dollars in a month, most of it spent at 3 AM when the product had almost no traffic and every one of those GPUs sat nearly idle.

This is the story that makes ML platform engineers care about scaling. A CPU web service that is over-provisioned wastes cents. A GPU service that is over-provisioned wastes a fortune, because the hardware is an order of magnitude more expensive and it does not get cheaper when it is idle. An H100 costs the same per hour whether it is running at 95 percent or 5 percent. That single asymmetry, a flat hourly cost paired with spiky demand, is the entire reason this lesson exists.

Idle GPU cost chart: a full day of hourly GPU utilization drawn as a green area that rises to a midday and evening peak and collapses overnight, against a red dashed provisioned-capacity ceiling billed twenty-four seven, with the wide red region between them labeled paid-for and idle, and callouts noting a flat $691 daily bill for eight cards at $3.60 an hour, an average utilization near 38 percent, and roughly $428 a day wasted on idle silicon

So the goal of this lesson is money, framed as engineering. You want to serve every real request fast enough, and pay for as little idle GPU as you can get away with. Everything that follows is one of two moves: push more requests through each expensive card, or run fewer cards when nobody is knocking. Keep those two levers in mind, because every technique in this lesson pulls on one of them.

Why GPUs, and the One Rule That Governs Everything

First, why a GPU at all. A CPU has a handful of powerful cores built to run one thing quickly. A GPU has thousands of small cores built to run the same operation across huge batches of numbers at once. A neural network is, underneath, a pile of matrix multiplications, which is exactly that kind of work. So a GPU can run a model tens to hundreds of times faster than a CPU. For anything bigger than a small model, GPU serving is not a luxury, it is the only way to hit a reasonable . But to scale a thing you have to know what it is made of, and two parts of a GPU set every limit that follows.

GPU package anatomy: a dark card illustration with high-bandwidth HBM3 memory stacks flanking a central die of streaming-multiprocessor tiles whose tensor cores are highlighted in green, an on-die L2 cache bar beneath them, and interconnect ports along the bottom showing NVLink at 900 GB/s card to card and PCIe Gen5 at 64 GB/s host to card, with legend cards explaining that VRAM is the hard ceiling on model size, the cores do the matrix math, and 3.35 TB/s of memory bandwidth feeds them

Now the one rule that governs all of GPU serving:

The model's weights must fit in the GPU's memory (VRAM), and they stay resident there for the life of the replica.

Which raises a question worth settling before you plan anything against it: how much memory does the card actually have? The product page is not the answer.

A recorded terminal session on a rented NVIDIA L40S. A comment reads that the spec sheet for this card says 48 GB, then asks the card. The command nvidia-smi with query-gpu for name, memory.total and memory.free in CSV format returns: NVIDIA L40S, 46068 MiB total, 45460 MiB free. The figure notes that 48 GiB is 49,152 MiB, so 3,084 MiB is gone before a single weight is loaded, and another 608 before anything of yours runs.
RecordedA rented L40S, asked directly. Cropped from a 91-column recording, not retyped

The card is sold as 48 GB. Forty eight gibibytes is 49,152 MiB, and this one reports 46,068. Three thousand and eighty four mebibytes are gone before you load anything, and a further 608 are taken by the driver's own context before a single tensor of yours arrives.

Two practical consequences, and the second is the one that bites.

Your usable budget is about six percent smaller than the number in the procurement spreadsheet. On a model sized to leave a couple of gigabytes of headroom on paper, six percent is the whole headroom, and the failure mode is not a warning at deploy time. It is an out-of-memory crash partway through the first request that needs a slightly longer context than the last one.

And the shortfall is not a rounding error you can wave away, it is structural. This class of card runs error-correcting memory, which reserves a fixed fraction of every chip, so it is present on every card of the type and it is not a defect. It is also not universal: cards using a different memory generation report their nominal size exactly, so a rule you learned on one card does not transfer to another.

The habit that follows is short. Never size a deployment from a product page. Run one command on the actual instance type you intend to rent, read what the card reports, and plan against that number.

You do not stream a model in from disk per request. You load it into VRAM once, when the replica boots, and it lives there. That makes VRAM the hard limit on what you can serve. Every card has a fixed amount, and the model plus its working memory has to fit inside it.

Here is the sizing math you will use constantly. Weights take the number of parameters times the bytes per parameter:

Model sizeFP32 (4 bytes)FP16 (2 bytes)INT8 (1 byte)
1B params~4 GB~2 GB~1 GB
7B params~28 GB~14 GB~7 GB
13B params~52 GB~26 GB~13 GB
70B params~280 GB~140 GB~70 GB

And that is only the weights. You also need room for activations and, for language models, the that grows with every token and every concurrent request. A 7B model in FP16 needs about 14 GB of weights, so on a 24 GB card you have roughly 10 GB left for everything else. Run out and the process dies with a CUDA out-of-memory error. This single constraint is why so much of GPU serving is about making the model smaller, and why the first question in any capacity discussion is always "does it fit."

Batching: The Cheapest Way to Make One GPU Do More

Before you buy more GPUs, you make each GPU do more work. A GPU running one request at a time leaves almost all of its thousands of cores idle, because the math for a single input barely fills them. The fix is dynamic batching: the server holds incoming requests for a tiny bounded window, gathers whatever arrives, and runs them all in a single forward pass. Because the card does the math in parallel, a batch of sixteen costs almost the same wall-clock time as a batch of one. This is the biggest cheap win in all of GPU serving, and you should turn it on before you touch anything else.

Batch size versus throughput and latency chart: a dual-axis curve where throughput in requests per second climbs steeply with batch size from 45 at batch 1 toward a flat ceiling near 740, while per-request latency stays low then rises sharply past a marked sweet spot around batch 16, with the point where big batches start costing real latency for little extra throughput called out as the knee

The chart shows why batching is free money right up until it is not. The first few doublings of batch size are nearly pure profit: many more requests served for almost no extra . Then you cross the knee, where a bigger batch takes so long to compute that per-request latency blows your budget while throughput barely improves. The craft is to pick the largest batch whose latency still fits your p99 target, then set a max batch size and a max wait window and let the server fire on whichever comes first.

Batching is only the first of several ways to get more out of one card. The others, roughly in order of return on effort, are: , storing the model's numbers in fewer bits so it both fits smaller and runs faster; distillation, training a small student model to imitate a big teacher so you can serve the student on a fraction of the hardware; and caching, putting a response cache in front of the GPUs so repeated inputs and shared prompt prefixes are answered in milliseconds without touching a card at all. A cache hit is the cheapest request there is, because it consumes zero GPU time.

Cost-lever taxonomy table summarizing the whole lesson on one card: seven levers, dynamic batching, quantization, distillation, response caching, MIG and time-slicing, autoscaling with scale-in, and spot for training, each with its typical throughput or cost multiplier, an effort badge from close to free to a real project, what it does, and when to reach for it, ordered top to bottom by return on effort

The order matters, and the multipliers stack. Batching and FP16 quantization are close to free and you should almost always do both, and together they can turn one card into the equivalent of three before you touch anything else. Distillation is a real project with training cost up front but permanent serving savings. depends entirely on how repetitive your traffic is. Stack the ones that apply to your workload and a single GPU can often do the work you thought needed three or four.

Telling Kubernetes to Give a Pod a GPU

None of this matters if the platform cannot even place your model on a GPU. On , a GPU is a countable resource, requested the same way you request CPU or memory. The catch is that a GPU cannot be fractionally shared by default: a pod either gets a whole card or it gets none. Here is a serving deployment that asks for one GPU, paired with an autoscaling policy driven by GPU load rather than CPU.

apiVersion: apps/v1
kind: Deployment
metadata:
  name: model-server
spec:
  replicas: 2
  template:
    spec:
      containers:
        - name: triton
          image: registry.internal/model-server:2026-07-11
          resources:
            limits:
              nvidia.com/gpu: 1        # one whole GPU per replica
              memory: "24Gi"
          readinessProbe:
            httpGet: { path: /v2/health/ready, port: 8000 }
            initialDelaySeconds: 120    # weights take time to load into VRAM
            periodSeconds: 5
---
# Scale on GPU load, not CPU. GPU work is invisible to CPU metrics.
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: model-server-scaler
spec:
  scaleTargetRef:
    name: model-server
  minReplicaCount: 1                    # a warm floor to hide cold starts
  maxReplicaCount: 12                   # ceiling set by GPU quota + budget
  cooldownPeriod: 300                   # wait before scaling down
  triggers:
    - type: prometheus
      metadata:
        query: avg(DCGM_FI_DEV_GPU_UTIL)  # real GPU utilization
        threshold: "70"

Read the important lines. The nvidia.com/gpu: 1 request is the whole game: it tells the scheduler this pod needs a physical GPU, so Kubernetes only places it on a node that has a free one. The initialDelaySeconds: 120 on the readiness probe exists because loading gigabytes of weights into VRAM is slow, and you do not want traffic routed to a replica that is still loading. And the autoscaler scales on DCGM_FI_DEV_GPU_UTIL, actual GPU utilization from NVIDIA's metrics exporter, because a fully loaded GPU can show almost no CPU usage. Scale on CPU here and you would never scale up at all, because the work that matters is invisible to the CPU-based signal a normal web autoscaler watches.

Sharing One Card: Stop Wasting a Whole H100

That "whole card or none" default is a problem the moment your model is small. A model that needs 7 GB still books all 80 GB of an H100 and every one of its cores, and you pay for the rest sitting idle. Modern NVIDIA cards give you two ways out, and they trade off differently.

GPU sharing panel with three rows on an 80 GB card: the top row shows one small model booking the whole card with about 91 percent hatched as idle but still billed, the middle row shows MIG carving the card into isolated hardware instances of 10, 10, 20, and 40 GB each with its own memory and cores, and the bottom row shows time-slicing where five small models take turns on the same card with high utilization but no memory isolation

MIG, short for Multi-Instance GPU, carves one physical card into as many as seven isolated smaller GPUs, each with its own slice of memory and cores. The isolation is real hardware isolation, so one tenant cannot starve another. Time-slicing is looser: several models take turns on the same card without hard memory isolation, which packs more of them on but lets a noisy neighbor hurt the others. Reach for MIG when tenants must not interfere, for example different teams or a -critical model that cannot tolerate a noisy neighbor. Reach for time-slicing when the models are friendly and you just want to squeeze more of them onto one card. Either way, the goal is the same: stop paying for 80 GB to run a 7 GB model.

NVIDIA built MIG precisely because large customers were wasting A100s and H100s on models too small to fill them. It is multi-tenancy solved in hardware, and on a platform serving many small models it is one of the highest-leverage cost tools you have.

Bin-Packing: The Scheduler Decides How Many Nodes You Rent

Sharing a single card is one half of the packing problem. The other half is how models are spread across nodes. GPU nodes are the bins and your models are the items, and how the scheduler packs them decides how many expensive nodes you keep switched on.

Bin-packing comparison: on the left, default scheduling spreads eight models thinly across four nodes so each node sits around 20 to 30 percent full yet fully billed, wasting the idle remainder; on the right, bin-packing scheduling fills two nodes to near capacity with the same eight models and releases the other two nodes entirely so they scale in and cost zero

Spread models thinly and every node is half-empty yet fully billed, so you pay for four cards to do the work of two. Pack them tight, filling each node before opening the next, and the same eight models fit on two nodes near capacity while the other two are released and stop billing entirely. Tight packing plus scale-in is one of the largest and least glamorous GPU savings there is. The trade is that a densely packed node has less headroom for a sudden spike, so you leave a little slack and pair packing with autoscaling that can add a node before the packed ones saturate. This is a scheduling decision, not a code change, which is what makes it easy to leave on the table and easy to win once you notice it.

Scale Out or Scale Up, and How to Split a Big Model

When one GPU is genuinely not enough, you can go in two directions, and they solve different problems. Horizontal scaling means more replicas of the same GPU behind a . This buys throughput and scales elastically with traffic, so you add replicas for the busy hours and drop them at night. What it does not do is make a single request faster, and it cannot help at all if the model does not fit on one card to begin with. Vertical scaling means a bigger GPU, from L4 to A100 to H100, or a multi-GPU node with the model split across cards. This is what lets a large model fit at all, and a stronger card can cut per-request latency, but it is lumpy and expensive and eventually you hit the biggest card that exists.

When a single model is too large for any card you can buy, you split it, and there are three distinct ways to spread the work across GPUs.

Three-way parallelism topology: data parallel copies the full model onto each of four GPUs and splits the inputs across them for throughput, tensor parallel cuts each layer into shards spread across all four cards so they exchange tensors every step, and pipeline parallel assigns whole consecutive layers to each card as pipeline stages that pass activations along

Data parallel copies the whole model onto each GPU and splits the inputs, which is really just horizontal scaling for serving: each replica handles a share of requests independently, but the model must fit on one card. Tensor parallel and pipeline parallel split the model itself, which is the only way to run a model too big for any single card. Tensor parallel cuts each layer across cards that must talk on every step; pipeline parallel gives each card a slice of the depth and passes activations along. For serving, data parallel is the workhorse and the other two are a necessity, not a preference: you reach for them only when the model will not fit, and then you pay for it in interconnect and complexity.

Interconnect is the word to underline, because splitting a model only pays off if the cards can talk fast enough.

NVLink versus PCIe bandwidth bar chart: NVLink 4 on H100 at 900 GB/s and NVLink 3 on A100 at 600 GB/s tower over PCIe Gen5 x16 at 64 GB/s and PCIe Gen4 x16 at 32 GB/s, with legend notes that NVLink is roughly fourteen times faster than PCIe and is required for a split model, while PCIe is fine for independent data-parallel replicas that barely talk to each other

NVLink, the direct card-to-card fabric inside a node, moves data more than ten times faster than the PCIe bus that connects a card to the host. A tensor-parallel layer exchanges tensors on every step, and NVLink keeps that traffic off the critical path where PCIe cannot. Two H100s with NVLink between them are a genuinely different machine from two H100s that can only reach each other over PCIe. If you plan to split a model across cards, the fabric between them is part of the spec you are paying for, and getting it wrong quietly caps the expensive silicon you bought.

Picking a Card: Sticker Price Is Not Cost Per Inference

The instinct when the bill is high is to reach for the cheapest card. That instinct is usually wrong, and understanding why is the difference between a serving bill that scales and one that surprises you. What matters is not the hourly price, it is dollars per inference: the hourly price divided by how many inferences that card actually completes in an hour.

Cost-per-inference bar chart across four GPU types running the same small model near each card's throughput knee: the T4 at $0.35 an hour and 200 requests per second costs about $0.486 per million inferences, the L4 at $0.75 an hour and 550 requests per second about $0.379, the A100 80GB at $3.60 an hour and 2100 requests per second about $0.476, and the H100 80GB at $4.50 an hour and 4800 requests per second the lowest at about $0.260, with a note that at 20 percent utilization every bar is five times higher

Run the numbers and the H100, the most expensive card by the hour, is the cheapest per request, because its scales faster than its price. But read the caveat, because it is the whole point: those numbers assume each card runs near its throughput knee. At 20 percent utilization every bar is five times higher, and a barely-used H100 becomes the single most expensive choice you could make, pairing the highest hourly price with the worst cost per inference. And VRAM still gates the decision before cost does. A model that needs 40 GB simply cannot run on a T4 or an L4, so you fit first and optimize cost per inference second.

The practical rule: pick the card by fit and by how busy you can keep it, not by hourly price. If a model fits on a small card and traffic is light, a T4 or L4 avoids paying for capacity you cannot fill. If traffic is heavy enough to keep a big card busy, the big card is cheaper per request. The expensive mistake is the idle premium card.

Autoscaling GPUs, the Cold-Start Tax, and Spot

Horizontal scaling is only cheap if the replica count actually tracks demand, which means autoscaling. But GPU autoscaling has a problem that CPU autoscaling does not: adding a replica is slow and expensive to warm up. A stateless web pod is ready in a second. A GPU replica has to find a scarce GPU node, pull a multi-gigabyte CUDA image, and load gigabytes of weights into VRAM before it can answer anything, commonly two to three minutes.

Autoscaling timeline over one day: a red step line of demand rising and falling across the hours, tracked by a blue dashed step line of running replicas that leads demand on the way up and lags it on the way down through a cooldown, with a green warm-floor band marking a minimum of two replicas that the pool never drops below, so cold starts are hidden and brief lulls do not trigger a costly re-warm

If you wait until the queue is already on fire to scale up, users feel the pain for the whole warm-up. So GPU autoscaling has a distinct playbook: keep a minimum warm floor of replicas so the next user is never greeted by a full cold start, scale up early on a leading signal such as headroom or queue depth rather than waiting for saturation, cache images on the node pool, and wait out a cooldown before scaling down so a brief lull does not cost you a cold start on the next spike.

Two more cost levers live around autoscaling. For training, which is fault-tolerant because it checkpoints, you run on spot or preemptible instances at sixty to ninety percent off, and you accept that the cloud can reclaim them mid-run. You almost never serve live traffic on spot, because a reclaimed serving node is a user-facing outage. Same fleet, opposite risk tolerance: aggressive spot for training, on-demand with a warm floor for serving.

Do the Math Yourself: A Sizing and Cost Calculator

Everything above rests on two calculations you should be able to do in your head: does the model fit in VRAM, and what does an inference cost once you account for . Here is both as code you can run and change. Adjust the parameter count, the precision, and the batching multiplier and watch what fits and what it costs.

The output makes the two rules concrete. Fit is a hard gate set by VRAM, so a card either qualifies or it does not. Among the cards that qualify, cost per inference is a throughput game, and the batching multiplier moves it more than the choice of card does. Set BATCH_GAIN back to 1 and every cost figure jumps, which is the numerical version of the earlier point: the batch window, not the hardware, is often the single biggest dial on the bill.

A Decision Map, and How the Heavy Users Keep the Bill Sane

Put it all together and you get an order of operations. When the bill arrives, work the cheap levers first and reserve buying hardware for last, because most GPU bills fall by half or more before you ever add a card.

GPU cost-optimization decision map: starting from the GPU bill is too high, a top-to-bottom sequence of questions each paired with an action, first turn on dynamic batching and FP16, next quantize further and then distill, then cache repeated inputs and prompt prefixes, then slice the card with MIG or time-slicing if it is too big for the model, then autoscale with a warm floor and send training to spot, and only at the very end scale out or up when there is a genuine capacity limit

None of this is theoretical. Every company serving models at scale has fought this exact fight. Character.AI wrote publicly about serving enormous volumes of language-model traffic affordably, and their toolkit is precisely this lesson: aggressive INT8 quantization, attention tricks that shrink the , and heavy caching of shared prompt prefixes so repeated context does not re-run on the GPU. They reported cutting serving cost per request by more than an order of magnitude, which is what made a free consumer chat product economically possible at all.

OpenAI, Anthropic, and other LLM providers lean on , where new requests join an in-flight batch token by token instead of waiting for a fixed window, keeping the GPUs saturated even under uneven traffic. It is dynamic batching taken to its logical extreme for autoregressive generation. Netflix, Uber, and DoorDash, serving thousands of smaller recommendation and ranking models rather than giant LLMs, focus on packing many models onto shared GPUs and on autoscaling down hard during off-peak hours, because at their model count, idle GPUs are the whole game.

The common thread across all of them is the same. Nobody solves a GPU cost problem by simply buying more GPUs. They make each card do more, through batching and , and they run fewer cards when demand is low, through sharing and autoscaling. That is the entire discipline of scaling GPU infrastructure, and it is why the number that ties this lesson together is requests per second per dollar. Take a card that costs two dollars an hour and naively serves 40 requests per second: that is 20 requests per second per dollar-hour. Turn on FP16 and dynamic batching and the same card can push 200 or more requests per second, five times the value on the exact same hardware. That multiplier, not the choice of card, is usually the difference between a serving bill that scales and one that gets you the note from finance.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

Why must a model's weights fit in the GPU's VRAM for serving?

Q2

Why should a GPU serving autoscaler scale on GPU utilization or queue depth rather than CPU usage?

Q3

What is the trade-off that dynamic batching makes?

Q4

A small model runs on an H100 that stays at about 15 percent utilization all day. Why is this often the most expensive choice per inference?

The honest answer for most platforms is both directions. You go vertical until the model fits and each request is fast enough, then you go horizontal to serve the volume, with autoscaling deciding how many replicas run right now.