There is no single best one, and no universal best size. The right choice depends on how your documents are written, how long one complete idea runs in them, and what your users ask. Short FAQ entries and long legal clauses do not want the same size.
The lesson's honest order is this. Start with fixed-size at 512 tokens and 50 overlap. Use structural or recursive splitting when your documents have real headings. Adopt semantic splitting only when a measurement on your own data shows it wins. Complexity you cannot justify with a measurement is just risk.
Size is also a cost dial. For a corpus of one million tokens with ten percent overlap, returning the top five chunks, 128-token chunks give about 8,700 embeddings and about 640 tokens per prompt. 1024-token chunks give about 1,090 embeddings but about 5,120 prompt tokens. Smaller chunks grow the index. Larger ones grow every prompt.
Two patterns ease the trade. Store metadata on every chunk: the document id, its position and its source. Position lets you fetch the neighbouring chunk, and it powers small-to-big retrieval, where you search small chunks for sharp matches and then hand the model the larger parent section. Anthropic's Contextual Retrieval adds a short generated sentence of context to each chunk before embedding. In their testing that step alone cut retrieval failures by roughly a third.
Whatever you pick, confirm it with RAG evaluation: a set of real questions with known source passages, checked after every change. The Chunking lesson walks through each strategy with diagrams and the full runnable chunker. Two other lessons measure chunking directly: splitting text into pieces and whether there is really a best chunk size.