Every vector database benchmark you will be shown is measuring something you do not care about. Queries per second with no recall number is not a capability, it is a tuning choice — any approximate index goes faster if you let it search less of the graph. Recall with no latency number is the same trick in reverse.
The only figure that means anything is recall at a fixed latency budget, on your own vectors, with your own filter patterns applied. All three qualifiers are load-bearing. Public datasets have different intrinsic dimensionality and cluster structure from your embeddings. Public benchmarks run without filters, and filters are where products genuinely differ. And they measure a freshly built, fully warm index, which is not what you will be running in six months.
The second thing worth saying before any product names: you may not need one of these. If your corpus is modest and Postgres already holds the data your filters run on, pgvector removes an entire system, keeps your joins, and deletes the sync pipeline between your source of truth and your index — where a large share of RAG staleness bugs live. Start there and add a dedicated engine when you can name the specific thing Postgres could not do.
Key takeaways
- Recall and latency are meaningless apart. Measure recall at k against an exact brute-force ground truth on your own vectors, at a fixed p95 latency and realistic concurrency.
- Filtered search is where these products genuinely differ. A selective filter can leave graph traversal stranded with no matching neighbours, and recall collapses silently on exactly the tenant-scoped queries you run most.
- Quantisation is a recall-for-memory trade you measure, never a free win. Rescoring top candidates against full-precision vectors recovers most of the loss.
- Namespace-per-tenant avoids the filtered-search problem structurally but multiplies per-index overhead; filter-per-tenant scales to many tenants and walks straight into it.
Recall at a fixed latency budget, on your data, is the whole measurement
Recall at k is the fraction of the true nearest k neighbours that the approximate index actually returned, where “true” means computed by exact brute-force scan over your own vectors. That ground-truth pass is expensive, you only run it once per dataset and embedding model, and without it you have no idea what your retrieval is missing.
Hold these constant across every candidate, or the comparison is noise:
- The embedding model and dimensionality. Change the model and the geometry changes, so every parameter you tuned is tuned for a different space.
- The dataset — yours, at production size. Ten thousand vectors tells you nothing about a hundred million, because the failure modes are graph structure and memory.
- k, and the filters you actually run. Measuring unfiltered recall and shipping a filtered application is the most common evaluation mistake in this category.
- Concurrency and warmth. A single-threaded query on a warm index is a different measurement from p95 under real concurrency with a partially cold cache.
Then report recall at k alongside p95 latency at that concurrency. One without the other is marketing.
Three more things belong in the same test and rarely appear in anyone’s benchmark. Index build and ingest time, because a rebuild you cannot finish in a maintenance window changes your architecture. Memory per million vectors at your dimensionality and index parameters, because that is the invoice. And recall after mutation — replay a week of writes and deletes, then re-measure. Graph indexes treat deletes as tombstones and degrade until rebuilt, so the number nobody publishes is recall on an index that has been alive a while.
Needs first-hand data: Build a ground-truth set with an exact brute-force scan over your production embeddings for 1,000 sampled queries, including the metadata filters those queries carry. Then measure recall at 10 and p95 latency for each candidate at your real concurrency, once on a fresh index and again after replaying a week of writes and deletes, and record memory per million vectors alongside it.
Index families and what each one actually costs
Flat, exact search. Brute force over every vector. Correct by definition, cost linear in corpus size. For a corpus in the low hundreds of thousands, or any query where a filter cuts hard, this is frequently right and it eliminates every tuning question in this article. Do not skip past it because it sounds unsophisticated.
HNSW. A multi-layer proximity graph traversed by greedy descent. Fast, high recall, the default in most products. What it costs: memory, because you store the graph as well as the vectors; rebuild expense, because construction is the slow part; and degradation under mutation, because deletes are tombstones. Three parameters matter — graph degree, build-time search width, query-time search width — and only the last is a runtime dial you tune per query class to trade latency for recall.
IVF and its variants. Partition the space into clusters, probe only the nearest few. Far cheaper in memory than HNSW and considerably more tuning-sensitive: probe too few cells and you miss neighbours just across a boundary. Worse, the partitioning is fitted to the data present at build time, so as your corpus drifts the clusters stop matching it and recall decays with no signal.
Disk-based approaches. Keep the graph on SSD with a compressed in-memory structure for routing, so corpora that will never fit in RAM become affordable. The trade is explicit: latency is bounded by random reads and sensitive to page cache behaviour. For a large mostly-cold corpus it is often the only economically sane option.
Quantisation. Scalar quantisation to eight bits is usually close to lossless for a large memory reduction. Product quantisation is more aggressive and definitely lossy. Binary quantisation is cheapest and unusable without a rescoring pass. What makes all of them work is the same pattern: search the quantised index wide, then rescore top candidates against full-precision vectors. Treat every quantisation setting as a recall change to measure, never a free memory win — that is the assumption made silently and discovered during an incident.
Filtered search is where products differ, and where recall collapses quietly
Production RAG almost never runs an unfiltered query. It filters by tenant, collection, date range, permission scope. So how a product combines filtering with approximate search matters more than its raw benchmark, and both strategies have sharp edges.
Post-filtering. Run the ANN search for the top k, then discard results that fail the filter. Trivially correct in ranking and catastrophic under selectivity: ask for 10, get 10 candidates, 9 fail the filter, return 1. Products compensate by over-fetching some multiple of k, a guess that gets worse as the filter gets more selective.
Pre-filtering. Restrict the search to the matching set. If that set is small enough to brute-force, this is exact and fast. If instead the engine traverses the ANN graph while skipping non-matching nodes, you hit the real failure in this category. The graph’s connectivity was built assuming all nodes are present. Skip 99% of them and greedy traversal reaches a node with no matching neighbours to hop to, terminates early, and returns whatever it found on the way. No error is raised. You get results. They are not the nearest neighbours.
And it fails precisely on the queries a multi-tenant application runs constantly: the selective ones. A tenant holding 0.1% of your vectors is the worst case, and that is every query that tenant makes.
What to do about it:
- Measure recall with your filters applied. This one step catches the problem before your users do.
- Below a selectivity threshold, brute-force the matching subset — faster and exact. Several products do this automatically; find out where the switch happens, because that threshold is now part of your architecture.
- For a low-cardinality dimension present in every query — tenant, region, collection — a separate index or namespace beats a filter, trading per-index overhead for exactness.
- Prefer engines that build filter awareness into graph construction, keeping same-tenant vectors connected, rather than treating the filter as post-processing.
Hybrid search and reranking, briefly
Dense vectors are bad at exact tokens. Part numbers, error codes, SKUs and rare acronyms are what users search for most confidently and what embeddings match least reliably. BM25 is excellent at exactly those and poor at paraphrase. On any real corpus, keyword and vector retrieval together beat either alone, and it is not close.
Combining them is mostly a ranking problem. Reciprocal rank fusion is the boring default that works, because it uses ranks rather than normalising two incompatible score scales. Weighted score fusion can do better once tuned and needs retuning whenever either retriever changes.
Then rerank. A cross-encoder reads the query and a candidate together rather than comparing two independent embeddings, which makes it far more accurate and far too slow to run over a corpus. So: retrieve wide with cheap methods, rerank narrow with an expensive one. For your database choice this matters concretely — native hybrid search and multi-phase ranking save you writing and maintaining fusion code.
Multitenancy: namespace per tenant or filter per tenant
This choice interacts directly with the filtered-search problem, which is why it belongs here rather than in an operations appendix.
Namespace per tenant. A separate index per tenant. No filter, so the recall collapse cannot happen, and deleting a tenant is dropping a namespace rather than a project. The cost is per-index overhead multiplied by tenant count — for tens of thousands of small tenants that becomes memory you cannot justify — plus cold-start latency for a tenant whose index is not resident.
Filter per tenant. One index with the tenant identifier as a filter. Cheap, scales to any number of tenants, and puts you in the filtered-search failure mode on every query. It also makes hard deletion a rewrite rather than a drop, which matters when a customer exercises a deletion right — see AI gateways for compliance for the surrounding obligations.
The honest answer for most multi-tenant products is a hybrid: namespaces for large tenants where the overhead amortises, one shared filtered index for the long tail, and brute-force search within a tenant’s subset when it is small — which for the long tail it almost always is. Whichever you choose, check whether the engine supports tenant-aware index construction, because that feature is what makes the filtered approach viable at all.
Needs first-hand data: For your own tenant size distribution, measure recall at 10 for your smallest, median and largest tenants under the filtered strategy, and compare against namespace-per-tenant for the same tenants with memory per tenant recorded in each configuration. The crossover point tells you where to draw the line between namespaces and the shared index.
pgvector

pgvector is a Postgres extension adding vector types, distance operators and both HNSW and IVFFlat indexes. It belongs first because for a large share of applications it is the correct answer, and the argument is architectural rather than about speed: your filters are real SQL predicates, your embeddings are written in the same transaction as the row they describe, and there is no pipeline between source of truth and index to go stale.
Pros
- Metadata filters are indexed SQL predicates, and the planner can filter first and scan exactly — sidestepping the ANN filtered-recall problem for selective filters
- Embeddings and their rows share a transaction, so there is no dual-write and no window where index and data disagree
- One system to back up, secure, monitor and grant access on, and joins mean results carry the rest of the record
Cons
- Index builds are memory-hungry, and a build exceeding your maintenance work memory falls back to a much slower path
- A retrieval workload competes with OLTP traffic for shared buffers and IO, so real volume wants a dedicated replica
- Scaling past one primary plus replicas means sharding, and hybrid search means combining with Postgres full-text search yourself
Best for: Teams whose corpus fits comfortably on Postgres hardware and whose filters run on data Postgres already holds, which is more teams than the market wants to admit.
Pricing: No licence cost; you pay for the Postgres you already run.
Pinecone

Pinecone is the fully managed option that exposes no index tuning surface at all, with an architecture separating storage from query compute so a large mostly-cold corpus is not priced like a hot one. Namespaces are a first-class primitive rather than a pattern you assemble, which makes per-tenant isolation the default path rather than a workaround for the filtering problem.
Pros
- Nothing to tune and nothing to operate, the correct trade for a team with no capacity to own ANN parameters
- Namespaces are first class, so tenant isolation does not depend on filtered-search behaviour
- Separating storage from query compute means idle data costs storage rather than provisioned memory
Cons
- Managed-only, so a residency, on-premise or air-gap requirement ends the evaluation immediately
- You cannot tune index parameters, which is fine until recall is unacceptable and there is no dial to reach for
- Usage-shaped metering is hard to forecast before you know your read and write mix at production scale
Best for: Product teams that want retrieval to work without anyone owning ANN tuning, with per-tenant namespaces available from the first week.
Pricing: Usage-based serverless metering on stored data plus read and write operations, with enterprise plans above.
Qdrant

Qdrant is an open-source Rust engine whose distinguishing feature is taking filtering seriously: payload indexes and filter-aware graph traversal rather than filtering applied afterwards. It also exposes quantisation explicitly, with rescoring, so the memory-versus-recall trade is a setting you measure rather than a hidden default.
Pros
- Filtering is designed into HNSW traversal with supporting payload indexes, the most direct answer here to the filtered-recall problem
- Scalar, product and binary quantisation with rescoring are exposed as configuration, so the memory trade is measurable
- Open source and self-hostable with the same engine behind its managed cloud, and multitenancy documented as a first-class concern
Cons
- Self-hosting means owning a stateful distributed system, including shard placement, replication and recovery
- The tuning surface is large — the price of the control, with defaults that are sensible rather than optimal for your data
- Hybrid search relies on configuring sparse vectors rather than a mature lexical engine underneath
Best for: Teams that need high recall under selective per-tenant filters and are willing to tune for it, self-hosted or managed.
Pricing: Open source with no licence cost, plus a managed cloud metered on cluster resources.
Weaviate

Weaviate is an open-source engine with a schema-first data model, native hybrid search, optional vectorizer modules that embed at ingest, and per-tenant shards including offloading of inactive tenants. That last feature addresses the long-tail tenant cost problem directly, which few engines do.
Pros
- Native hybrid search with built-in fusion, so combining lexical and dense retrieval is not code you maintain
- Per-tenant shards with offloading of inactive tenants, the specific answer to namespace overhead across many small tenants
- Schema-first model makes filters, data types and cross-references explicit rather than conventional
Cons
- The schema and module model is opinionated and a genuine learning investment before anything works
- Letting the database own embedding couples retrieval quality to its module versions, and changing model becomes a full re-index
- Memory footprint at scale needs active attention, and per-tenant shard counts have practical ceilings
Best for: Teams that want hybrid search and per-tenant shards working out of the box rather than assembling BM25, fusion and tenancy themselves.
Pricing: Open source with no licence cost plus a managed cloud metered on stored vector volume and compute.
Milvus

Milvus has the widest index selection here — HNSW, IVF variants, disk-based and quantised — and a distributed architecture separating query, index and data nodes so read and write scaling are independent. It is built for the scale where a memory-resident graph index stops being affordable.
Pros
- The broadest index selection in this list, so you match the index family to your memory budget rather than accepting one implementation
- Separated query, index and data nodes let you scale ingest and search independently under continuous update
- Fully open source with a large ecosystem and no licence ceiling on scale
Cons
- Many components and external dependencies, so self-hosting is a platform project rather than a deployment
- Index flexibility is only an advantage if someone owns choosing and tuning, and most teams do not have that person
- At small scale you carry all the complexity and none of the benefit
Best for: Teams with a very large corpus and a platform team able to own a distributed stateful system, or who intend to run it managed from the start.
Pricing: Open source with no licence cost; the cost is the infrastructure it runs on and the engineers who operate it.
Zilliz

Zilliz is the commercial company behind Milvus, marketing a managed “vector lakebase” built on that engine. Treat these two as one choice with two deployment modes rather than as competitors: the query semantics share a lineage, so a self-hosted proof of concept and a managed production deployment are not two different products.
Pros
- Managed Milvus keeps the index flexibility while the distributed-systems operations stop being your problem
- Same engine lineage means a self-hosted evaluation transfers directly to the managed deployment
- The company stewarding the engine also runs the service, so roadmap and managed features are aligned
Cons
- The commercial offering is managed, so residency or air-gap requirements push you back to self-hosted Milvus and its operational load
- One company stewards both engine and service, concentrating roadmap and pricing risk in a single relationship
- Forecasting cost requires understanding your read, write and index-type mix in advance
Best for: Teams that want Milvus’s index flexibility at large scale without staffing the distributed-systems work themselves.
Pricing: Managed usage-based tiers on compute and storage with enterprise agreements above; the underlying engine remains open source at no licence cost.
Chroma

Chroma is the developer-first option: an in-process library with a small API that makes the first week of a retrieval prototype nearly frictionless, with a server mode to grow into. Its role in a serious architecture is usually the thing you build the prototype on before deciding what production needs.
Pros
- Starts embedded with no server, which removes a service from a prototype entirely
- A small API surface, so nobody stalls on schema design before writing retrieval code
- The same API works against a server, so outgrowing embedded mode is not a rewrite of your call sites
Cons
- Not the choice for large-scale production retrieval, and teams that let the prototype become the architecture find out late
- Filtering and tenancy controls are thinner, so the failure modes in this article are less mitigated
- Operational maturity, observability and scaling controls trail the established engines
Best for: Prototypes, local development and genuinely small corpora where embedded mode removes a service from the architecture.
Pricing: Open source with no licence cost, plus a managed hosted offering metered on usage.
MongoDB Atlas Vector Search

Atlas Vector Search puts vector indexes next to the documents they describe and exposes retrieval through the aggregation pipeline your application already uses. It is the pgvector argument in a document model: no separate index to synchronise and no second source of truth.
Pros
- Vectors live beside their documents, so there is no sync pipeline between source of truth and index
- Filtering, projection and post-processing use the aggregation pipeline rather than a new query language
- One operational surface — scaling, backup, access control — for application data and retrieval together
Cons
- This is an Atlas capability, so the argument largely does not transfer to a self-managed deployment
- The tuning surface is deliberately narrow, so recall under selective filters is something you measure and live with
- Vector workloads compete with operational traffic unless you separate search nodes, and index variety is limited
Best for: Teams already on Atlas whose documents are the source of truth and whose filters run on fields those documents already carry.
Pricing: Part of the Atlas platform, metered on cluster and search-node resources rather than as a separate per-vector charge.
OpenSearch

OpenSearch is a mature lexical search engine with vector search added, which gives it the strongest hybrid-search story among the general-purpose options: BM25, filters, aggregations and access control are already there. Multiple k-NN backends give you quantised and disk-based options for the memory question.
Pros
- A mature lexical engine underneath means hybrid search is native rather than fusion code you own
- Filters, aggregations, access control and index lifecycle management already exist, and in production those matter more than raw ANN speed
- Self-hostable including air-gapped, with managed offerings available from cloud providers
Cons
- JVM cluster operations — shard sizing, heap tuning, lifecycle policies — become somebody’s permanent job
- Vector index memory pressure interacts badly with a shared cluster, so vector workloads generally need dedicated nodes
- An enormous configuration surface whose defaults are not tuned for vector recall
Best for: Teams already running this stack for lexical search who want hybrid retrieval without introducing a second datastore.
Pricing: Open source with no licence cost; managed offerings meter cluster resources and storage.
Vespa

Vespa is a search and ranking engine where multi-phase ranking is a first-class concept rather than something you assemble in application code. Filters, lexical matching, vector search and machine-learned ranking models all evaluate in one query next to the data — the architecture the reranking section argues for, expressed in an engine.
Pros
- Multi-phase ranking is native, so retrieve-wide-then-rerank-narrow runs inside the engine rather than across two services
- Combines filters, lexical matching, vector search and ranking models in a single query next to the data
- Built for large corpora with real-time updates, where bulk-rebuild architectures struggle
Cons
- The steepest learning curve here by a wide margin; application packages, schemas and ranking expressions are a specialism
- Operationally heavy, and a small team will spend more on understanding it than it saves them
- Overkill for a modest RAG corpus, where the flexibility is complexity you pay for and never use
Best for: Large-scale retrieval and ranking workloads with real-time updates, where multi-phase ranking inside the engine justifies a specialist owner.
Pricing: Open source with no licence cost, plus a managed cloud metered on the resources your application consumes.
turbopuffer

turbopuffer is object-storage-first by design, offering vector and full-text search as a serverless service where namespaces are the intended tenancy primitive. That targets a specific and common shape: a multi-tenant application with a long tail of mostly-idle tenants, where one large hot index is the wrong economics.
Pros
- Object storage as the primary tier makes cold, large corpora dramatically cheaper than memory-resident indexes
- Namespace-per-tenant is the intended pattern rather than a filter, avoiding the filtered-recall problem structurally
- Vector and full-text search in one service, and idle tenants cost storage rather than provisioned compute
Cons
- Cold-query latency is fundamentally higher than an in-memory index, and caching is what closes the gap — so latency depends on access patterns
- Managed service only, so residency and air-gap requirements are a blocker
- A newer entrant, so there is less accumulated operational knowledge when something behaves oddly
Best for: Multi-tenant applications with many mostly-idle tenants, where per-tenant namespaces on cheap storage beat one large hot index.
Pricing: Usage-based metering on stored data plus queries and writes, reflecting the object-storage-first design rather than provisioned cluster size.
How to choose
Run this in order. Steps one and two eliminate most of the list for most teams.
1. Ask whether you need a separate system at all. If your corpus fits on Postgres hardware and your filters run on relational data you already store, pgvector or Atlas Vector Search removes a system, a sync pipeline and a class of staleness bug.
2. Write down your filter patterns before looking at any benchmark. Tenant, collection, date, permission scope, and the selectivity of each. If your dominant pattern is a selective tenant filter, filtered-search behaviour and namespace support are your primary criteria and raw QPS is close to irrelevant.
3. Compute a ground truth on your own vectors. Exact nearest neighbours for a thousand sampled queries with their real filters. Everything after this is measured against that set.
4. Pick a latency budget first, then compare recall inside it. A budget forces you to include concurrency and cache warmth.
5. Measure memory per million vectors and recall after mutation. These decide your bill and your six-month experience, and neither appears in a vendor benchmark.
6. Decide tenancy explicitly. Namespaces for large tenants, shared filtered index plus brute force for the long tail, and a documented crossover point.
| Product | Deployment shape | Picks itself when |
|---|---|---|
| pgvector | Postgres extension | Your corpus fits and your filters are already SQL |
| Pinecone | Managed serverless | Nobody should be tuning an index, and namespaces matter from day one |
| Qdrant | Open source or managed | Selective per-tenant filters at high recall, with someone willing to tune |
| Weaviate | Open source or managed | You want native hybrid search and per-tenant shards without assembling them |
| Milvus | Self-hosted open source | Very large corpus and a platform team to own a distributed system |
| Zilliz | Managed Milvus | You want Milvus’s index range without operating it |
| Chroma | Embedded or hosted | Prototype, local development, or a genuinely small corpus |
| MongoDB Atlas Vector Search | Managed platform feature | Atlas already holds the documents and the filter fields |
| OpenSearch | Self-hosted or managed | You already run it for lexical search and want hybrid retrieval |
| Vespa | Self-hosted or managed | Multi-phase ranking inside the engine is worth a specialist |
| turbopuffer | Managed serverless | Many mostly-idle tenants, each in its own namespace |
Frequently asked questions
Do I need a vector database if I already have Postgres?
Often not. pgvector keeps embeddings in the same transaction as your data, lets filters be indexed SQL predicates, and deletes the synchronisation pipeline that causes most RAG staleness bugs. Move to a dedicated engine when you can name the specific limit you hit — build time, memory, sharding, or a filtered-recall requirement Postgres cannot meet.
What recall should I be targeting?
That is a product question, not a database question. Measure how answer quality degrades as recall drops on your own evaluation set: some applications are indistinguishable at 0.9, others break because the one document that mattered ranked eleventh. Decide the threshold from downstream answer quality, then buy the latency that gets you there.
Why did recall drop when I added a metadata filter?
Because approximate graph traversal assumed all nodes were present. With most filtered out, greedy search reaches a point with no matching neighbours and stops early. Fix it by brute-forcing small candidate sets, using a namespace instead of a filter for low-cardinality dimensions, or choosing an engine that builds filter awareness into the index.
How much memory will a million vectors need?
Measure it — the answer is a function of your dimensionality, index parameters, graph degree and quantisation setting, all of which are yours to choose. Build the index on a representative million-vector sample at your real dimensionality, read the number off, then multiply. Any figure quoted without your dimensionality attached describes somebody else’s corpus.
Related reading
- Best AI gateways — the hub, and where retrieval sits relative to the rest of the LLM request path.
- Best RAG frameworks — parsing, chunking and reranking, which affect answer quality more than index choice does.
- Best semantic caching tools — the same ANN machinery used for caching, and why false hits are the risk there.
- Best LLM evaluation tools — retrieval-level metrics, without which you cannot tell retrieval failures from generation failures.
- Best AI gateways for compliance — per-tenant deletion and residency requirements that shape your tenancy design.