"Can I cancel my order?" and "Can I cancel my subscription?" are almost the same question with completely different answers.
If your app is caching its LLM answers by meaning, that pair is how it hands a customer a confident, wrong one.
Response caching is the cheapest speedup there is: store the answer keyed to the question, then serve it straight back next time instead of paying for the call again.
The interesting part is deciding what counts as the same question.
Exact match is a plain lookup: the incoming text has to be identical, character for character. Instant, never wrong, and it misses nearly everything, because people don't ask twice the same way.
Semantic caching turns each question into an embedding (a list of numbers standing in for its meaning) and hunts for a stored one whose numbers sit close by. Now "what's the refund policy?" reuses the answer you wrote for "can I get my money back?"
Here is the bit that changed how I think about it: close is a threshold you pick. Loosen it and you catch more paraphrases, and also start answering questions nobody asked, like the cancel pair up top. Tighten it and the cache is safe and barely fires. Finding that number on your own traffic is the actual work.
Not to be confused with prompt caching, which Anthropic and OpenAI run on their end: it saves the processing of a long instruction block you resend every call, so that part is billed at a fraction of the rate. The model still runs and still writes a fresh answer, so it makes the call cheaper where response caching removes it.
Don't make the model answer a question it has already answered. Just be sure it is the same question.
Quick check before you scroll: What's the core difference between exact-match caching and semantic caching for LLM responses?
Full breakdown + the answer: frankduah.me/learnings/2026-09-18-caching-llm-responses-the-cheapest-speedup
New here? I post a bite-size AI / ML concept like this every day. Follow me for the daily drop, and it compounds fast. Why I do it: https://lnkd.in/gK8knHDH
#SemanticCaching #LLMOps #AIEngineering #AI #LLM #AIAgents #MachineLearning
The answer
Exact-match caching only returns a cached response when the new input is identical to a past one. Semantic caching embeds the input's meaning and matches it against past inputs by similarity, so it also catches reworded or paraphrased questions.