Training happens once. After that, the tokenizer only uses its list of merges. To cut a new word, it starts from single characters or bytes. Then it applies the learned merges in the order it learned them, earliest first, until no neighbouring pair is on the list.
Order matters. Each later merge was learned on text where the earlier merges were already glued. Replaying them in the same order means new text is cut the same way training saw it. That is also why the same word is always cut the same way. The tokenizer does not guess.
Real tokenizers add a few things on top of the lesson's teaching version. They start from bytes, not characters, so any text in any language, or any emoji, can always be cut and nothing is ever unknown. They train on far more text, including many languages and code. They run far more merges. They also split text into words with a more careful rule than spaces; the GPT-4 and GPT-4o tokenizers keep digits in groups of up to three.
OpenAI's tiktoken README sums up what BPE gives. It is reversible and lossless, so tokens turn back into the exact text. It works on any text. It makes text shorter. And it lets the model see common word parts like "ing".