Real embeddings have hundreds of numbers, so here is the same arithmetic with three. These numbers are made up. Only the method is real. The question is (2, 1, 0). Document A is (4, 2, 1). Document B is (0, 1, 3).
First, the dot product: multiply numbers in the same position, then add. Question and A give 10. Question and B give 1. Next, the length of each list: the square root of the sum of its squares. That is about 2.236 for the question, 4.583 for A and 3.162 for B.
Then divide the dot product by the two lengths multiplied together. A scores about 0.976. B scores about 0.141. A points almost the same way as the question, so A ranks first.
That is all a semantic search does. It repeats these steps for every document and sorts. The lesson also ran a small real search over ten lesson titles. The question was "How can I stop one user from sending too many requests to my API?", asked in English, Spanish and Hindi. "Rate Limiting" came first every time, at 0.749, 0.729 and 0.736. The Hindi question was never translated. The model alone linked it to the English lesson.