A semantic cache helps most when traffic repeats and answers are stable: support desks, documentation assistants and internal knowledge bots, which get the same questions in many wordings. Prompt caching helps any workload that resends a long, stable prefix, which is why agents benefit so much.
A semantic cache needs care in three cases. If the answer depends on who is asking, such as their account or region, key the cache on the user plus the question, or do not cache it. If the answer changes over time, such as a price, give entries a short TTL, a time to live. If almost every question is new, the hit rate stays near zero and the lookups are pure overhead.
Last, tie the cache to your documents. When you update a document, answers already in the cache were written from the old version. Expire or version the cache together with the index, or it will keep serving last week's policy. This is the usual cache invalidation problem, with higher stakes.
The RAG Caching and Cost lesson has two calculators to run with your own numbers. One prints the monthly bill with and without a cache. The other shows how a loose threshold serves wrong answers. It is the best next step before you put a semantic cache in front of real users.