Paying Once for the Same Question
An AI system that recognizes a question it has already answered, asked in other words, can reuse the answer instead of paying the model again. The cost per question falls as questions repeat.
An AI system that recognizes a question it has already answered, in other words, can reuse the answer instead of paying the model again. The cost per question then falls as questions repeat.
Most AI systems pay the model again for every question, including one they answered an hour ago. A semantic cache stores each answer with a numerical fingerprint of what the question meant, and when a new question means the same thing, it returns the stored answer for the price of a lookup.
Matching Meaning Instead of Words
An ordinary cache matches identical requests only, and people rarely ask the same thing in the same words.
"What are your hours?" is answered and stored. "When do you close?" misses, and the model is called again: one question, paid for twice.
"What are your hours?" is answered and stored. "When do you close?" means the same thing, so the stored answer comes back. One model call serves both.
Questions with close meanings get close fingerprints, and when their similarity passes a threshold (on a scale where 1.0 is identical) the stored answer is served. The threshold is the design decision: set it low and a question gets an answer meant for another one; set it high and only near-identical questions reuse an answer.
Two of our four engines use this pattern. WatchBase, in build, reuses a narration at 0.92 or above. AskBase, available now, sets the bar higher where a wrong answer costs more: 0.95 for report summaries and 0.97 for answers from a firm’s documents. An answer from a firm’s documents is checked against the asking person’s permissions before it is served and expires after a week; a report summary expires after 30 days.
How the Cost Curve Changes
Without a cache, the model bill grows in step with use. With one, repeated questions stop reaching the model, so the bill grows with new questions rather than with all of them.
Illustrative: an FAQ-style workload, not a measured hit rate. The cache answers 20% of requests in month one, 40% by month three, 60% by month six and 75% by month twelve.
More people send more questions, and every question is a model call. The thousandth person asking about refunds costs what the first did.
Questions already answered come from the cache. The thousandth person asking about refunds costs a lookup, and the model bill grows with new questions, not with every question.
A hit rate belongs to your traffic rather than to the cache, so an honest figure names its workload: “58% on this one”, not “58%”.
Where It Pays
The return is highest where similar requests repeat at volume.
Where It Does Not Pay
A cache helps only when requests repeat. A question nobody has asked in any form is a miss: the model answers it and the answer is stored for next time. Tailored output, such as a proposal or a one-off email, has no reusable answer, and reasoning across several documents or a long conversation depends on context too particular to match.
The risk is a threshold set too loose, which serves a stored answer to a question that only sounds similar. Set the bar highest where a wrong answer is costly, make a cached answer pass the same checks as a fresh one, and keep the repetitive request paths apart from the new ones: the cache handles the volume, and the model handles what is new.
Which questions is your business paying an AI to answer twice?
Book Free Assessment