The short answerA GPU is a general parallel chip, sold by NVIDIA, AMD and others, that runs almost any AI code anywhere; a TPU is Google's own chip built only for the matrix maths inside neural networks and used through Google Cloud, so pick a GPU for flexibility and the widest software support, and a TPU for large, steady training or serving jobs on Google Cloud.
This page is the free part.
The course goes deeper on GPU and TPU
₹999 in India/$49 everywhere elseonce, for the whole course
The AI Engineering course covers GPU and TPU across a run of lessons, not one page. These 3 alone are about 180 minutes of step-by-step reading, every one with code you run in the browser, all with a quiz.
Why GPU serving gets expensive and how to fix it: the VRAM ceiling, batching, quantization, distillation, caching, card sharing, bin-packing, parallelism, autoscaling, spot, and horizontal vs vertical scaling.
Why LLM inference is slow and expensive, and the techniques that fix it: the KV cache, continuous batching, PagedAttention, quantization, and speculative decoding.
Thousands of small cores run the same step on many numbers at once.
Runs PyTorch, TensorFlow, JAX and custom code (CUDA).
You can buy one, rent one from any cloud, or have one in a laptop.
vs
Tensor Processing Unit
TPU
one huge multiplication grid
Google's custom chip (an ASIC) made only for machine learning.
Its core is a 128 x 128 grid that multiplies matrices as data flows through.
Runs JAX, TensorFlow and PyTorch through the XLA compiler.
Used through Google Cloud: Compute Engine, GKE and Google's AI platform.
Many workers vs one multiplication gridA GPU has many flexible cores that each fetch their own data. A TPU's matrix unit is a fixed grid where data flows from cell to cell. The grid reuses every number many times, which saves trips to memory, but it can only do one kind of job well.
GPU vs TPU, side by side
Read across a row to compare one thing. Every word that may be new is explained just below the table.
GPU compared with TPU
Aspect
GPU
TPU
Who makes it
NVIDIA, AMD, Intel, Apple (inside its own chips).
Google only.
Built for
Any highly parallel maths: graphics, AI, science.
Neural network maths, mostly matrix multiplies.
How the maths runs
Thousands of cores, plus tensor cores for matrix maths.
A 128 x 128 systolic array per matrix unit (MXU).
Software
PyTorch, TensorFlow, JAX, plus CUDA and many custom kernels.
JAX, TensorFlow, PyTorch through XLA. Custom ops are harder.
Where you get one
Buy it, rent it from any cloud, or use one in a laptop.
Google Cloud: Compute Engine, GKE, Google's AI platform.
Peak BF16 per chip (vendor figure)
H100 SXM: 1,979 TFLOPS, which NVIDIA marks "with sparsity".
v5e: 16 GB. v6e: 32 GB. TPU7x: 192 GB at about 7.4 TB/s.
Scaling to many chips
NVLink inside a server (900 GB/s on H100 SXM), networks between servers.
Chips wired in a torus; a TPU7x pod has 9,216 chips.
Weak spots
Power and price; supply can be tight.
Dynamic shapes, heavy branching, custom ops, high precision (Google's own list).
Best for
Research, mixed workloads, anything off Google Cloud.
Large, steady training and serving on Google Cloud.
Words on this page, in plain English
Chip
A processor: the part of a computer that does the maths.
CPU
Central processing unit. A few powerful cores that are good at many different tasks, one after another.
Core
One worker inside a chip that can run instructions.
Matrix multiply
Multiplying two grids of numbers. Almost all the work in a neural network is this one operation.
ASIC
Application-specific integrated circuit: a chip designed for one kind of job. A TPU is an ASIC for neural networks.
Systolic array
A grid of tiny multiply-and-add cells. Numbers pulse through it like blood through a heart, so each value is reused many times without going back to memory.
HBM
High-bandwidth memory: very fast memory stacked next to the chip. How much a chip has decides how big a model fits.
TFLOPS
Trillions of floating point operations per second. A measure of raw maths speed.
BF16
A 16-bit number format used in AI. It is less precise than normal 32-bit numbers but twice as cheap to move and multiply.
CUDA
NVIDIA's software platform for programming its GPUs. A lot of AI code is written for it.
XLA
A compiler that turns model code into instructions for a chip. TPUs need it.
When to use GPU, when to use TPU
Real situations, and the pick we would make in each one.
Pick GPU
You are learning AI on a laptop or a single cloud machine.
Every tutorial and library works on GPUs. Even a laptop GPU, like the one in our test, runs small models.
Pick GPU
Research code with new layers you wrote yourself.
Google says models with many custom operations suit GPUs better. TPUs want standard matrix maths.
Pick TPU
Training a large model for weeks with big batches on Google Cloud.
This is the exact case Google lists for TPUs: matrix-heavy, large batches, long runs.
Pick TPU
A recommendation model with huge embedding tables.
Google lists ultra-large embeddings as a TPU strength; newer TPUs add SparseCores for them.
Pick GPU
Serving a chat model to users from AWS or Azure.
TPUs are a Google Cloud product. On other clouds, GPUs are the accelerator you can rent.
Pick Both
A model written in JAX that you want to run cheaply at scale.
JAX runs on both through XLA. Price the same job on each before you commit.
Picking a chip, using Google's own rulesMost questions lead to a GPU. The path to a TPU is narrow but real: big, regular, long jobs on Google Cloud.
we ran this, here is what happened
Hands-on: A matrix multiply on this Mac's GPU vs its CPU
We do not have a TPU. You can only use one through Google Cloud, and we did not want to show you numbers we cannot reproduce on this page. So the TPU side below uses Google's published figures, clearly labelled.
What we can measure is the idea behind both chips: parallel hardware beats a general CPU at matrix maths, but only when there is enough work to share out. We multiplied two square matrices of numbers on the CPU and then on the GPU of an Apple M4 laptop, using PyTorch.
For a size n, one multiply costs 2 x n x n x n operations. We timed 10 runs after 3 warm-ups, and repeated the whole script five times. The numbers below are the medians.
Where it ran: Apple M4 (10-core GPU), macOS 15.6, PyTorch 2.14.1 with the MPS backend (Apple's GPU interface), float32. Five full runs on 4 October 2026.
def bench(device: str, n: int, dtype: torch.dtype) -> dict:
g = torch.Generator().manual_seed(0)
a = torch.randn(n, n, generator=g).to(device=device, dtype=dtype)
b = torch.randn(n, n, generator=g).to(device=device, dtype=dtype)
def sync() -> None:
if device == "mps":
torch.mps.synchronize()
for _ in range(3):
a @ b
sync()
times = []
for _ in range(RUNS):
t0 = time.perf_counter()
a @ b
sync()
times.append(time.perf_counter() - t0)
s = statistics.median(times)
return {"n": n, "ms": round(s * 1000, 2), "tflops": round(2 * n**3 / s / 1e12, 3)}
From scripts/labs/compare/gpu_matmul.py. torch.mps.synchronize() waits for the GPU to finish; without it the timer would only measure the moment the work was handed over.
real output
size CPU (ms) GPU (ms) GPU speed-up
512 0.41 0.88 0.5x (the CPU wins)
1024 4.57 2.03 2.3x
2048 75.48 8.66 8.7x
4096 287.31 152.17 1.9x
8192 2383.89 1384.00 1.7x
accuracy vs float64 (1024 x 1024):
cpu_fp32_max_err 0.000207 gpu_fp32_max_err 0.000207 gpu_fp16_max_err 0.0702
Medians of five runs, from out-gpu-summary.json (summarize_gpu.py reads the five out-gpu-matmul-run*.json files). The table layout is ours; every number is copied from that file.
The results
What we measured
GPU (this Mac)
CPU (this Mac)
Small job, 512 x 512too little work: the CPU wins
0.88 ms
0.41 ms
Medium job, 2048 x 2048the GPU is 8.7 times faster
8.66 ms
75.48 ms
Big job, 8192 x 81921.7 times faster on a busy, shared GPU
1.38 s
2.38 s
Best speed reached (float32)a TPU7x chip is listed at 2,307 TFLOPS in BF16 (vendor figure)
about 2 TFLOPS
about 0.65 TFLOPS
GPU speed-up by matrix size, measuredBelow about 1000 x 1000 the CPU is quicker. In the middle the GPU is up to 8.7 times faster. On our busy, shared GPU the gap narrowed again at the biggest sizes.
What this shows
Parallel hardware only pays when the job is big enough to split. At 512 x 512 the CPU finished first, because handing work to the GPU has a fixed cost. At 2048 x 2048 the GPU was almost 9 times faster. A TPU pushes the same idea further: it gives up flexibility to spend nearly all of its silicon on one giant multiply grid.
What this test does not show: This is a laptop GPU, not a data-centre GPU or a TPU, so compare the shape of the results, not the sizes. During the test the same GPU was also running a local language model in the background and the machine was busy (load average 17 to 28), which is why the larger sizes dip. Both chips gave correct answers: against a float64 reference, the biggest error was 0.0002 in float32. The script is scripts/labs/compare/gpu_matmul.py in our repository.
Common mistakes
Comparing peak numbers straight across vendors.
NVIDIA's H100 BF16 figure is quoted with sparsity; Google's TPU figures are not marked that way. Peaks are also rarely reached. Compare the time and cost of your own job.
"A TPU is just a faster GPU."
It is a different design. It is excellent at big, regular matrix maths and weaker at branching, custom ops and changing shapes, as Google's own docs say.
Timing GPU code without waiting for it.
GPU calls return before the work is done. Call torch.mps.synchronize() or torch.cuda.synchronize() before you stop the clock, as our script does.
Sending tiny jobs to an accelerator.
In our test the CPU beat the GPU at 512 x 512. Batch small requests together so each trip to the chip carries enough work.
Ignoring memory.
A model that does not fit in the chip's memory will not run, however fast the chip is. Check HBM per chip first: 16 GB on a v5e, 80 GB on an H100 SXM.
Questions people ask
Which is best, GPU or TPU?
Neither in general. GPUs are more flexible, run everywhere and have the most software. TPUs can be a better deal for large, matrix-heavy jobs on Google Cloud. Google's own guidance sends models with many custom operations to GPUs and long, large-batch training to TPUs.
Does ChatGPT use GPU or TPU?
OpenAI has used NVIDIA GPUs in Microsoft Azure. Microsoft wrote in March 2023 that it linked tens of thousands of NVIDIA GPUs for OpenAI, and quoted OpenAI's president saying this made ChatGPT possible. Neither company publishes a full list of chips in use today.
Can TPUs replace GPUs?
For some jobs, yes: Google lists large, matrix-heavy training and serving as what TPUs are built for. For others, no: Google's docs list workloads TPUs are not suited to, such as programs with lots of branching, high-precision maths, or custom operations in the training loop.
Is Google TPU as good as NVIDIA GPU?
On paper the newest chips are close. Google lists TPU7x (Ironwood) at 2,307 TFLOPS in BF16 with 192 GB of memory; NVIDIA lists the H100 SXM at 1,979 TFLOPS in BF16 with sparsity and 80 GB. Those are different generations measured differently, so test your own model before you choose.
Do I need a GPU or TPU to learn AI engineering?
No. Most learning work, like calling models, building RAG and running small models locally, runs on a normal laptop. Our matrix test ran on a laptop GPU. Rent bigger chips only when a real job needs them.
Lessons that go deeper
From the AI Engineering course, in the order we would read them.
GPU vs TPU matters once you serve real models. Our AI Engineering course has 203 lessons on RAG, evals, agents, serving and MLOps, many built around a real experiment like the one on this page. 10 lessons are free to read, with no card needed.
the hands-on parts are real runs, like this one
course 2
AI Engineering
Take models from notebook to production, with labs on real models.