Top-k and top-p exist to stop that drift. Both are filters. A filter removes some tokens before the draw, so they cannot be picked at all.
Top-k keeps only the k most likely tokens, where k is a whole number such as 5 or 40. With top_k 5 after "The capital of France is", the five kept tokens add up to 0.7347. After rescaling, Paris gets about 0.85 of the draw. The other 151,931 tokens can no longer win.
Top-p keeps the most likely tokens in order until their odds add up to at least p, a number between 0 and 1. With top_p 0.5, Paris alone holds 0.6219, which is more than half. So Paris is the only token kept, and it is drawn every time. In the lab, all 40 runs of the fact prompt began with "Paris.".
The key difference: top-k always keeps the same number of tokens, whether the model is sure or not. Top-p adapts. It keeps one token when the model is sure and many when it is not. That is why top_p 0.5 cut the fact prompt to 26 different texts but left all 40 story texts different.
Note that top_k 1 keeps only the top token, so it behaves exactly like temperature 0. The lab got 1 text out of 40 for every prompt.