Semantic caching gets adopted because someone looked at an inference bill and someone else said hit rate. The number that decides whether it was a good idea is a different one, and almost nobody measures it: the rate at which the cache returns a confidently wrong answer to a question it should have missed on.
A cache miss costs money and a second of latency. A false hit costs correctness, raises no error, writes no warning, and is invisible from inside the system — the user gets a fluent, plausible answer to a question nobody asked. That asymmetry is the whole subject. It should decide your similarity threshold, which surfaces you enable it on, and in a good number of cases whether you enable it at all.
There is also a cheaper thing most teams skip on the way here. Exact-match caching on a complete deterministic key costs nothing in correctness and catches retries, refreshes, duplicated agent steps and replayed evaluation runs, which together are a real share of traffic in almost every LLM application. Semantic caching is the second step. Teams reach for embeddings first because it sounds more interesting, then discover their exact-duplicate rate was material and unmeasured.
So: the mechanics of where semantic caching goes wrong and how to bound it, then the tools — split honestly between caching primitives you assemble and caching as a feature you switch on in a gateway.
Key takeaways
- The failure mode is a false hit, not a miss. Negation, entity swaps, date-relative questions and unit changes all produce prompt pairs that are semantically close and have different correct answers.
- Exact-match caching on a complete request key — model, parameters, system prompt, tools, full message list — has zero correctness cost and should ship first.
- A cache shared across tenants is a data-leak mechanism the moment any prompt carries customer context, and per-tenant namespaces destroy the hit rate that made semantic caching attractive.
- A cached RAG answer built on a document that has since changed is now wrong, and nothing in the cache can tell.
How semantic caching works, and the four ways it returns a wrong answer
The mechanism is short. Embed the incoming prompt, search an approximate-nearest-neighbour index of previously cached prompt embeddings, and if the nearest neighbour’s distance is under a threshold, return the completion stored against it instead of calling the model.
The assumption buried in that sentence is that semantic similarity implies answer equivalence. It does not, and no embedding model was trained on that objective. Embeddings put topically related text near each other. “These two questions have the same correct answer” is a different relation and nobody optimised for it.
Four failure classes follow directly, and each produces a near-perfect similarity score:
Negation. “Which customers renewed last quarter” and “which customers did not renew last quarter” differ by one token carrying almost no embedding weight. The correct answers are complements. This is the fastest way to demonstrate the problem to a sceptical colleague.
Entity and parameter swaps. “Revenue for Q1” versus “Q2”. “Order 44718” versus “44719”. “Berlin” versus “Bern”. Identifiers and proper nouns are a small fraction of the vector, and swapping one changes the answer completely.
Date-relative language. “How many signups this month” is the same string today, tomorrow and next month. The prompt is stable, the correct answer is not, and similarity search happily matches the string to itself across the boundary.
Units, currencies and directions. Dollars versus euros, kilos versus pounds, “cheapest” versus “most expensive”. Superlatives invert on a word that barely moves the vector.
Now the part that decides the feature. Tighten the threshold enough to eliminate those failures and the hit rate collapses toward what exact-match caching would have given you anyway. Loosen it for a hit rate worth reporting and you serve wrong answers with no signal. A cosine similarity of 0.95 says two questions are about the same topic. It does not say their answers are interchangeable, and treating it as if it does is the actual bug.
Where semantic caching genuinely pays: high-volume, low-variance, read-only, non-personalised surfaces. Documentation and FAQ assistants where thousands of people ask the same fifteen questions in different words. Classification of similar inputs. Summary generation over a stable corpus.
Where I would not enable it: anything parameterised by an identifier or a number, anything phrased relative to now, anything carrying per-customer context, and anything inside an agent loop — a subtly wrong intermediate in step three of a ten-step run compounds into an outcome nobody traces back to a cache.
One control makes a real difference and most setups omit it: a second-stage verifier on candidate hits. Let the ANN search propose a neighbour, then check before serving — a cross-encoder, a small model asked whether the two questions have the same answer, or a deterministic rule that both prompts contain the same numbers, dates and named entities. That last one is cheap, boring, and catches three of the four classes above.
Needs first-hand data: Take 500 real prompt pairs from your own traffic that an embedding model scores above 0.9 similarity and label by hand whether each pair has the same correct answer. Plot false-hit rate and hit rate against threshold from that labelled set. That curve, on your own prompt distribution, is the only honest basis for choosing a threshold — and it is the artifact no vendor can give you.
Do exact-match caching first, with a complete key
The deterministic version has no correctness cost when the key is complete, and completeness is the whole job.
The key must cover everything that can change the response: model identifier including version, temperature, top-p, seed, maximum tokens, stop sequences, response format or schema, the full tool definitions, the system prompt, and every message in order. Serialise canonically — sorted keys, stable number formatting — and hash it.
Then go hunting for the things that break it, because they are usually already in your code:
- A current timestamp or date interpolated into the system prompt on every request. This defeats exact-match caching entirely and is the most common cause of a zero hit rate.
- A request, trace or session ID inside the prompt rather than in metadata.
- A prompt assembled from a structure with nondeterministic ordering, so one logical request serialises two ways.
- Randomised few-shot example selection nobody remembers adding.
Fixing those is worth more than any cache product, and it is worth doing before you evaluate one.
What exact-match catches is larger than people expect: client retries after a timeout that already succeeded server-side, browser refreshes, users resubmitting, duplicated steps in agent loops re-reading the same context, evaluation suites replaying cases in CI, and fan-out patterns where several code paths make the identical classification call.
One distinction to keep straight. Provider-side prompt caching — a discount on repeated prefix tokens — is a different mechanism from your response cache, and they compose rather than compete. Prefix caching saves prefill compute on a call you still make; a response cache skips the call. Do not read a provider’s prompt-cache discount as evidence that you need no application cache, and do not report the two as one saving.
If temperature is above zero, exact-match caching does change behaviour: identical requests return identical responses instead of varied samples. That is a product decision rather than a correctness bug, but it is a decision — a creative surface that suddenly stops varying generates support tickets.
Needs first-hand data: Log a canonical hash of every outbound request for a week — model, parameters, system prompt, tools, full message list — and compute the exact-duplicate rate per surface and per tenant. Then recompute it with volatile fields (timestamps, request IDs) excluded from the hash. The gap between those two numbers is free saving you are currently throwing away, and the absolute rate tells you whether semantic caching is worth any of the risk that follows.
Cache scope is a tenancy decision, and getting it wrong is a data leak
A cache keyed on prompt text and nothing else is shared across every user of your application. That is the default in most setups, including ones assembled from good primitives.
Consider what that means with a semantic layer on top. Tenant A asks about their own contract, and the answer lands in the cache carrying their data. Tenant B asks a semantically close question, the ANN search finds A’s entry, the distance is under threshold, and B receives A’s answer verbatim. No exploit, no bug report, no log line that looks wrong. This is not an exotic edge case — it is what the architecture does when nobody scopes it.
The fix is namespacing, and the strong version is physical rather than logical: the tenant identifier in the key namespace, and a separate index namespace per tenant so an ANN search cannot return someone else’s neighbour. Applying a tenant filter after retrieval is weaker for a reason worth understanding — filtered approximate search degrades in ways that are hard to see, which the vector database guide covers. For a cache, “the filter mostly works” is not a standard anyone should accept.
Other dimensions belong in the key too, or they cause wrong answers rather than leaks: the acting user where responses are personalised, the plan tier when the system prompt differs by tier, locale, and any flag that changes the prompt.
Then the awkward truth nobody puts in the pitch. Per-tenant scoping destroys the hit rate that made semantic caching attractive. The hit rate came from many different people asking similar things. Scope it per tenant and every tenant warms its own cache from cold. For a product with a long tail of small accounts, most tenants never accumulate enough traffic to reach a useful hit rate, and your headline saving turns out to have come almost entirely from your three largest customers.
There is a legitimate exception: a shared namespace for prompts that provably carry no tenant context — documentation questions, generic how-tos, “what does this error mean”. Building it requires reliably classifying which prompts are which, and the failure direction matters. Default to private and promote to shared only on an explicit allowlist of prompt templates. Defaulting to shared unless proven private is precisely how the leak above happens.
Invalidation, TTL, and the RAG problem nobody solves
TTL is your only real invalidation mechanism unless you build more. Choose it from how fast the underlying data changes, not from what makes the hit rate look good — those pressures point in opposite directions and only one of them has a dashboard.
Event-driven invalidation is harder here than on a normal cache. With exact keys you can tag each entry with the entities it touched and purge by tag when a record changes. With semantic entries you do not know which entities a natural-language question referenced unless you extracted them at write time. So extract entity tags when you store the entry if you want purge-ability, or be honest that TTL is your entire invalidation story.
RAG is the case that breaks cleanly. A cached answer was generated from a specific set of retrieved chunks. Those documents get updated, corrected or deleted. The cached answer is now wrong, it still matches incoming prompts perfectly, and nothing in the cache has any way to know.
Three responses, in descending order of rigour:
- Content version in the key. Hash the retrieved chunk identifiers and their revisions into the key. Correct, and it destroys most of the saving, because you must run retrieval before you can compute the key — and retrieval is often a meaningful share of the cost you were avoiding.
- A global cache epoch bumped on every index rebuild. Coarse, cheap, honest. Fine if you reindex nightly, unusable if you reindex continuously.
- A short TTL sized to your indexing cadence. The pragmatic default, worth documenting explicitly so the next engineer knows the staleness window was a decision.
Two smaller things. Do not cache errors, rate-limit responses or truncated completions — a cached 429 is a self-inflicted outage. And cached responses are available in full immediately, so most implementations replay them through the same token animation to keep the interface consistent.
Finally, log every cache decision with the similarity score, the matched prompt and the key namespace. Without the matched prompt you cannot investigate a false-hit complaint, because the wrong answer is fluent and the user’s report will be “it said something about the wrong customer”. That log is also where you discover the threshold is too loose — see LLM observability tools for where those events should land.
GPTCache

GPTCache is the open-source project that defined this category, and it is a primitive rather than a product: a Python library composing an embedding function, a vector store, a similarity evaluator and an eviction policy into a cache you wire into your own call path. Separating the similarity evaluator from the raw ANN distance is its most valuable design decision, because that is exactly where a second-stage verifier belongs.
Pros
- Explicit control over embedding model, vector store, evaluator and eviction, which is what honest threshold tuning requires
- The pluggable evaluator is the natural place for a cross-encoder or entity-matching check on candidate hits
- Runs in-process against a store you already operate, so no prompt text reaches a caching vendor
Cons
- A library, so tenancy namespacing, invalidation, metrics and eviction are code you write and must get right
- Python-only in practice, so a polyglot estate fronts it as an internal service or duplicates the logic
- Defaults will produce every false-hit class above, and nothing in the library warns you that it is happening
Best for: Python teams with one high-traffic surface worth optimising properly, who want to own the embedding model, the threshold and the verification step rather than accept a gateway’s single knob.
Pricing: Open source with no licence cost; you pay for the embedding calls it makes and the vector store you point it at.
Redis

Redis is where most exact-match caches already live, and with vector similarity search it can also serve as the ANN index behind a semantic cache. The argument is consolidation: the deterministic layer that should ship first and the semantic layer that may ship later run on one system you already operate and already know how to reason about.
Pros
- Most teams already run it, so exact-match caching is a config change rather than a new on-call surface
- Vector search in the same system means the semantic layer introduces no second datastore
- TTL, eviction and key namespacing are mature, and those mechanics matter more for correctness than the ANN part
Cons
- You are still building the semantic cache: embedding calls, threshold, verification and invalidation are your code
- Memory-resident indexes make a large semantic cache expensive per gigabyte next to disk-based stores
- Mixing vector and ordinary cache workloads on one instance hides the capacity problem until it bites
Best for: Teams already running Redis who want complete exact-match caching this week and the option to add a semantic layer later without adopting another datastore.
Pricing: Free-to-run editions alongside a commercial managed cloud and enterprise licensing; the managed meter is provisioned memory and throughput rather than per request.
Upstash

Upstash is a managed Redis-compatible store with HTTP access and per-request metering, plus a managed vector product alongside it. It appears in a caching article rather than a general infrastructure one because of serverless: a function that lives for 200 milliseconds cannot maintain a connection pool, and HTTP access removes that problem.
Pros
- HTTP access works from serverless functions and edge runtimes where a persistent client is impractical
- Per-request metering fits spiky, low-baseline traffic where provisioned memory would idle most of the day
- A managed vector product under the same account, so the semantic layer is not a second vendor relationship
Cons
- Per-request pricing inverts against you at high steady volume, where provisioned capacity wins
- Still a primitive: thresholds, tenancy scoping, verification and invalidation are your application’s problem
- HTTP adds a round trip relative to an in-VPC cache, paid on every miss as well as every hit
Best for: Serverless and edge deployments where connection-per-invocation makes a traditional client painful and traffic is spiky enough that provisioned memory would idle.
Pricing: Usage-based per request with storage and bandwidth metered separately, plus fixed-capacity plans for steady workloads.
Portkey

Portkey is now Prisma AIRS AI Gateway, part of Palo Alto Networks, and it represents the other shape here: caching as a gateway feature enabled per route rather than a library you assemble. It supports a deterministic mode and a semantic mode, and because caching sits in the same control plane as routing, keys and cost tracking, hit rate and spend appear in one place.
Pros
- Caching becomes a gateway config change, applying to every service routing through it with no application code
- Both modes with per-route configuration, so semantic runs only where it is safe and deterministic runs everywhere else
- Hit rate is visible next to cost and latency, which is how you notice whether the feature earns its risk
Cons
- Now part of Palo Alto Networks, so independent-vendor assumptions about pricing and roadmap no longer hold
- The threshold is a knob rather than a pipeline, with no clean place for the second-stage verifier that cuts false hits
- The gateway sees a request, not your domain model, so per-tenant scoping depends on metadata you thread correctly — and the failure is a leak, not an error
Best for: Platform teams who already want a gateway for routing, key management and cost tracking, and would rather switch caching on per route than build it in every service.
Pricing: Managed usage-based tiers metered on gateway traffic with enterprise agreements above; the meter is not itemised per feature.
Helicone
![]()
Helicone is an LLM observability product shaped as a proxy, with caching as one of the features that comes from sitting in the request path. Adoption is a base URL change, the lowest-friction integration here. It has announced that it is joining Mintlify, which is worth factoring in where caching is meant to be load-bearing.
Pros
- Integration is a base URL change, so caching and request logging arrive together rather than as two projects
- Cache hits appear in the same request log as everything else, so the saving is observable without extra instrumentation
- Open source with a self-host path, so prompts and cached responses can stay inside your infrastructure
Cons
- It has announced it is joining Mintlify, so confirm caching stays first-class rather than merely maintained
- Caching is one feature of an observability product, so expect fewer controls over similarity, verification and eviction
- A proxy in front of every call adds a hop and an availability dependency whose timeout behaviour you must verify
Best for: Teams whose primary need is LLM observability and who will take caching as a switch on the proxy they were adopting anyway.
Pricing: Open source and self-hostable at infrastructure cost, with a managed tier metered on logged requests and retention.
LiteLLM

LiteLLM straddles both shapes: a Python SDK you call in-process and a standalone proxy you deploy, with caching in either mode across several backends including a semantic option. Because its main job is normalising provider APIs, the cache key is computed against a normalised request, which makes deterministic caching coherent across providers rather than per-SDK.
Pros
- Adoptable as a library in one service or as a shared self-hosted proxy, so caching follows your existing architecture
- Multiple cache backends behind one interface, so moving from in-memory to Redis is configuration
- Open source and self-hostable, so no prompt or cached response leaves your network
Cons
- Self-hosted means you operate a proxy on the critical path, with scaling, upgrades and the pager attached
- Caching is one feature on a very large surface, with no natural hook for second-stage verification
- Broad provider support means parameter normalisation edge cases, and a key that omits a parameter returns a wrong answer rather than an error
Best for: Teams already using LiteLLM to normalise providers who want caching in the same layer without adding a vendor to the request path.
Pricing: Open source with no licence cost for the SDK and proxy, plus a commercial enterprise tier; you pay for the cache backend you run.
Cloudflare AI Gateway

Cloudflare AI Gateway is caching as an edge platform feature: point provider calls through it and you get caching, rate limiting, retries and request logging with nothing to deploy. It is the only option here where a hit is also a network-latency win, because the response comes from a location near the user rather than from your region.
Pros
- Hits are served at the edge, so the saving is latency as well as tokens
- No infrastructure to run, and it sits in front of whichever provider you call rather than one model vendor
- Bundles rate limiting, retries and request logs into the same layer
Cons
- Requests and responses transit a third party, which is a data-flow review and sometimes a hard blocker
- Cache behaviour is the platform’s model rather than a pipeline, so there is nowhere to add a stricter check
- Per-tenant separation depends on cache-key controls you configure correctly, with a cross-tenant leak as the failure mode
Best for: Teams already on Cloudflare’s platform who want caching, retries and rate limiting in front of provider APIs without operating anything.
Pricing: Usage-based within the platform’s AI offering, metered on requests proxied and log retention rather than cached bytes.
How to choose
The order matters more than the shortlist, because steps one and two frequently make the rest unnecessary.
1. Measure before you cache. Log a normalised hash of every request for a week and compute your true exact-duplicate rate per surface. This number decides whether any of this is worth engineering time.
2. Ship exact-match caching with a complete key. Then fix the timestamp in your system prompt and re-measure. This is often most of the available saving at zero correctness cost.
3. Gate semantic caching per surface, not per application. A surface qualifies only if it is read-only, carries no per-tenant data, contains no time-relative language and takes no numeric or identifier parameters. Most surfaces fail at least one.
4. If you enable it, add a verifier and a log. Entity, number and date matching between the incoming and matched prompts catches most of the damage cheaply.
5. Choose gateway or primitive on who owns the code. Many services needing coverage without touching each one: gateway. One high-traffic surface needing real tuning: primitive.
6. Set TTL from data freshness and handle RAG explicitly. Content version in the key, a global epoch on reindex, or a documented staleness window. Choosing none is choosing the third by accident.
| Tool | Shape | Exact and semantic | Picks itself when |
|---|---|---|---|
| GPTCache | Primitive you assemble | Both, fully configurable | You need to own the threshold and add verification |
| Redis | Primitive you already run | Exact natively, semantic via vector search | Consolidating both layers onto one system |
| Upstash | Managed primitive | Exact, semantic via its vector product | Serverless or edge, where connections are the problem |
| Portkey | Gateway feature | Both, per route | You want coverage across services without shipping app code |
| Helicone | Proxy feature | Caching alongside observability | Observability is the primary need |
| LiteLLM | Library or self-hosted proxy | Both, multiple backends | You already normalise providers through it |
| Cloudflare AI Gateway | Edge platform feature | Platform-managed caching | You are on that edge and want the latency win too |
Frequently asked questions
How much will semantic caching actually save?
Nobody can tell you without your prompt distribution, and a vendor’s headline hit rate was measured on somebody else’s traffic. Measure your exact-duplicate rate, ship deterministic caching, and treat the remaining semantic upside as something you have to prove. On genuinely repetitive surfaces the saving is real; on parameterised or per-tenant surfaces it is usually small and risky.
What similarity threshold should I use?
There is no portable number, because it depends on your embedding model and prompt distribution. Build a labelled set of high-similarity prompt pairs from your own traffic, mark which pairs share a correct answer, and read the threshold off that curve. Any threshold chosen without that set was chosen by feel.
Is provider prompt caching the same as semantic caching?
No. Provider prompt caching discounts repeated prefix tokens on a call you still make. A response cache skips the call. They compose, and reporting them as one saving makes your cost model wrong in both directions.
Can I share a semantic cache across tenants?
Not safely, unless every prompt in the shared namespace provably contains no tenant data. The safe pattern is per-tenant namespaces by default with an explicit allowlist of generic prompt templates promoted to a shared namespace. Defaulting to shared and filtering afterwards is how one customer’s answer reaches another.
Related reading
- Best AI gateways — the hub, and where caching sits among the other gateway responsibilities.
- Best LLM cost tracking tools — how to attribute the saving so you can tell whether caching worked.
- Best vector databases — the ANN index underneath any semantic cache, and why filtered search degrades quietly.
- Best LLM observability tools — where cache hit and miss events should land alongside completions.
- Best OpenAI-compatible proxies — the layer most teams already have, and the natural home for a deterministic cache.