A language model writes one token at a time. A token is a small piece of text. Each new token depends on all the tokens before it, so the model looks back at the whole history on every step.
Inside the model, attention turns each token into two vectors, a key and a value. Once computed, they never change, because the token behind them never changes. So servers compute them once and keep them in GPU memory, also called VRAM. That store is the KV cache. Without it, every step would redo the work of all earlier steps, and generation would get slower and slower as the reply grew.
The cost is memory. The KV cache grows with every token, for every request running at the same time. The lesson gives an illustration: for a model like Llama-2 13B the cache is on the order of 1 megabyte per token. At 4,096 tokens of context, that is about 4 GB per user. On an 80 GB card holding a 13 GB model, four such users already need more memory for their caches than for the model.
So the number of users one GPU can serve is mostly a question of how well you use memory for the KV cache. That is the problem PagedAttention solves.