Buyer’s Guide

Best RAG Frameworks

Written by Govind Kumar Lohar. Reviewed for technical accuracy by Deepak Gupta and Bhaskar Suthar on · Review panel

  • rag
  • llm
  • infrastructure
  • open-source

Independent buyer’s guide. No vendor paid to be included, ranked or described a particular way. Written for engineers, architects and the people who sign off on their tooling budget. Editorial policy.

The most common RAG bug I see reported as a model problem is a table that got flattened into prose during PDF parsing. The numbers are still in the text, stripped of their column headers, in an order nobody can reconstruct. Retrieval finds the chunk. The model reads it and produces a confident number that is wrong. Then the team spends two weeks on the prompt.

That is the shape of most RAG work. Answer quality is determined by the data pipeline — parsing, chunking, metadata, retrieval strategy, reranking — and almost none of it by which framework you imported. Frameworks matter for a different reason: they decide how much of your control flow you hand over, how hard the system is to debug, and how expensive it is to leave.

So there are two separate questions here. Where does retrieval quality come from, and how much framework are you willing to accept in exchange for the connectors? A useful signal on the first: LlamaIndex, arguably the company that knew this category best, has repositioned around document parsing, OCR and agent workflows rather than being “the RAG framework”. That is not a marketing accident. It is where the hard part actually was.

Key takeaways

  • Quality comes from parsing, then structure-aware chunking, then filterable metadata, then hybrid retrieval, then reranking. A table flattened into prose is unrecoverable downstream, and no framework fixes it.
  • Frameworks divide into libraries you call for one step and orchestration layers that own your control flow. The second kind is expensive to leave and hard to debug — test debuggability before adopting one.
  • Without a retrieval-level metric you cannot tell a retrieval failure from a generation failure, so teams tune the prompt when the chunker was the problem.
  • Unstructured is an ingestion and parsing platform, not an orchestration framework, and DSPy is an optimiser rather than a pipeline. Comparing them as if they were interchangeable produces a bad decision.

Where RAG quality actually comes from, in order

Work these in sequence. Doing them out of order is how teams spend a quarter and end up where they started.

1. Parsing. A PDF table flattened into prose is unrecoverable. The column headers are gone, the row boundaries are gone, and a merged cell has either duplicated or dropped its value. No chunker, no reranker and no frontier model reconstructs that, because the information was destroyed before retrieval existed. The same class of loss comes from multi-column layouts read in the wrong order, headers and footers interleaved into body text, scanned pages with no OCR pass, formulas rendered as gibberish, and text that only exists inside an image.

The diagnostic is simple and almost nobody runs it: print the extracted text for twenty representative documents and read it. That is the highest-value hour in any RAG project. What you want from a parser is structure preserved — tables emitted as tables, headings as headings, reading order intact — not a wall of characters with the layout thrown away.

2. Chunking that respects document structure. Fixed character counts split sentences mid-clause, separate a table from its caption, and cut a procedure in half so neither piece is actionable. Structure-aware chunking follows headings and element boundaries, keeps a table with its header row, keeps a list intact, and prepends a breadcrumb of parent headings so each chunk is interpretable on its own.

Chunk size is a real trade rather than a best practice. Small chunks embed precisely and lose the context that makes them useful; large chunks carry context and dilute the embedding until it matches everything vaguely. Fixed overlap is a crude fix that inflates your index. Parent-document retrieval — embed small chunks, return the enclosing section — usually beats tuning the size, and it is a few lines of code rather than a framework feature.

3. Metadata you can filter on. Source, document type, effective date, version, owning team, permission scope. Extract it at ingest, because retrofitting metadata means reprocessing the entire corpus. A large share of “the assistant gave outdated information” complaints are a missing effective-date filter, not a model failure, and a large share of the rest are a missing permission filter, which is a security bug rather than a quality one.

4. Hybrid retrieval. Dense vectors are weak on exact tokens: part numbers, error codes, SKUs, rare acronyms — the things users search for most confidently. Keyword search is excellent at those and poor at paraphrase. Combining them beats either alone on real corpora, and it is not close. The vector database guide covers the fusion mechanics and which engines do it natively.

5. Reranking. A cross-encoder reads the query and a candidate together instead of comparing two independently computed embeddings, so it is far more accurate and far too slow to run over a corpus. Retrieve fifty to a hundred candidates cheaply, rerank, keep the top few. After fixing parsing, this is the largest quality gain per hour of effort available to you.

Note what is missing from that list: the prompt. Prompt work matters, but it is the cheapest thing to try and produces the most visible change, which is exactly why teams reach for it first and stay there. Fix parsing before chunk size, fix chunking before retrieval strategy, and touch the prompt last.

Needs first-hand data: Take twenty documents that represent your corpus’s worst cases — a scanned invoice, a multi-column report, a spreadsheet exported to PDF, a contract with nested tables — and run each through two or three parsers. Record, per document, whether tables survived as tables, whether reading order held, and how many numeric values were mangled. That table decides your parser, and it is the artifact no vendor benchmark will give you.

How much framework are you willing to accept

There are two kinds of thing being sold as a RAG framework, and conflating them is the most expensive mistake in this category.

A library you call for one step. A parser, a chunker, a reranker client, a vector store adapter, a set of loaders. You call it, it returns data, your code decides what happens next. Leaving costs you one import and a function signature.

An orchestration layer that owns your control flow. It decides the sequence, holds the state, handles retries, and calls back into your code. Your application’s logic now lives inside its abstractions and your stack traces run through its scheduler. Leaving means rewriting the flow.

The second kind is genuinely worth it for some problems and routinely adopted for problems that do not need it. Before you take one on, run three tests.

The debuggability test. A ten-step run returns a bad answer. Can you see the retrieval query as issued, the full candidate list with scores, the reranked list, and the exact prompt that reached the model with the assembled context in it — or do you get the final output and a trace of function names? If it is the latter, you will debug by print statement, and you will do it during an incident.

The isolation test. Can you call the retriever on its own, outside the pipeline, from a notebook? Iteration speed on retrieval quality is the thing that determines whether you actually improve it, and a retriever you can only invoke through a full run makes every experiment expensive.

The upgrade test. How much did the abstractions change across the last few releases? If your control flow is theirs, every upgrade is a migration, and abandonware-adjacent version pinning is how teams end up on a two-year-old release with known bugs.

My own judgment: take the libraries and connectors generously, because dozens of loaders and adapters is real work you should not repeat. Adopt an orchestration layer only when your control flow is genuinely complex — branching, multiple cooperating agents, human-in-the-loop approval, durable resumption after a crash — and you would otherwise be writing a scheduler yourself. A straightforward retrieve, rerank, generate pipeline is around fifty lines of your own code, completely inspectable, and I would write it every time. Plain function calls are underrated, and they never break on upgrade.

Evaluation is not optional, because otherwise you cannot attribute the failure

A bad answer has exactly two possible causes, and they need different fixes: the right chunk was never retrieved, or the right chunk was retrieved and the model did not answer from it faithfully. Without a retrieval-level metric you cannot tell which, so you guess, and guessing looks like prompt tuning.

Retrieval metrics. Recall at k against a labelled set: real user questions, each paired with the chunk that should answer it. Building that set is the work, and it is not much work — fifty to a hundred hand-labelled examples drawn from actual user questions is worth more than a thousand synthetic ones, because the synthetic set was generated from the documents you already retrieve well.

Generation metrics. Faithfulness, meaning every claim in the answer is supported by the retrieved context, and answer relevance. These only mean something once you know retrieval succeeded, which is why the order matters.

Then log the assembled prompt with the retrieved chunk identifiers and their scores for every production request. That single log line is what turns “the assistant said something wrong about our refund policy” from an unfalsifiable report into a five-minute investigation. The LLM evaluation tools roundup covers the harnesses, and LLM observability covers where those traces should land.

Needs first-hand data: Label 100 real user questions with the chunk that should answer each, then measure recall at 5 and at 20 for three configurations: your current pipeline, the same pipeline with hybrid retrieval, and hybrid plus a reranker. Record the answer-quality change alongside it. That comparison tells you where your remaining budget should go, and it usually says parsing rather than prompting.

LlamaIndex

LlamaIndex homepage

LlamaIndex has repositioned around document parsing and OCR plus agent workflows rather than being the general-purpose RAG framework it was known as. That move is the most useful piece of information in this article: the team with the most exposure to real RAG projects concluded that the hard, valuable part is getting documents into usable structured form. You can still use the library pieces individually, and you can take the parsing products without adopting anything else.

Pros

  • Its parsing and OCR products target the failure that actually determines RAG quality, which is a more useful thing to buy than orchestration
  • Parsers that preserve tables and reading order remove the unrecoverable loss at the top of the pipeline
  • The library half remains usable a piece at a time — loaders, node parsers, retrievers — without adopting a runtime
  • An agent workflow layer is there when your control flow genuinely needs one

Cons

  • The repositioning means a lot of existing tutorials describe abstractions that have moved, and following a stale guide costs a day
  • Managed parsing sends your documents to a third party, which is a data-flow review and sometimes a hard blocker
  • The abstraction surface is large and has changed repeatedly, so upgrades are not free on long-lived code

Best for: Teams whose retrieval quality is limited by document parsing — tables, scans, complex layouts — rather than by orchestration.

Pricing: Open-source libraries at no licence cost, with the managed parsing and platform services metered on documents or pages processed.

LangChain

LangChain homepage

LangChain is one company shipping three things worth distinguishing: LangChain the framework, LangGraph the agent runtime, and LangSmith the commercial observability and evaluation platform. Its real asset is integration breadth — the loaders, vector store adapters and model wrappers you would otherwise write. LangGraph is the part to take seriously if your flow genuinely branches, because graph-shaped control flow with explicit state is a better answer than a chain of implicit calls.

Pros

  • The largest set of loaders, store adapters and model integrations anywhere, and that is genuine work you do not have to repeat
  • LangGraph expresses control flow as an explicit graph with state, which is the right shape for branching and resumable runs
  • LangSmith traces show each step, each retrieval and each prompt as sent, which satisfies the debuggability test directly
  • A very large community, so almost any integration question has already been asked and answered somewhere

Cons

  • The abstraction surface is enormous and has churned, so code written against an older version needs revisiting rather than upgrading
  • When the framework owns control flow, stack traces run through it and reasoning about failures gets harder as the graph grows
  • LangSmith is a commercial hosted product, not an open-source component, so the visibility you need has a vendor and a bill attached

Best for: Teams that want the connector breadth and have genuinely graph-shaped control flow, and accept a commercial platform for the tracing that makes it debuggable.

Pricing: The framework and agent runtime are open source at no licence cost; LangSmith is a commercial platform metered on traces ingested and users, with enterprise tiers above.

Haystack

Haystack homepage

Haystack, from deepset, models a pipeline as an explicit graph of components with declared inputs and outputs. That design choice pays off in exactly the places the previous section cares about: what runs is inspectable rather than inferred, individual components can be exercised alone, and the pipeline definition serialises into a reviewable artifact instead of being scattered through application code.

Pros

  • Explicit component graphs with declared contracts, so the pipeline is readable without tracing through implicit call chains
  • Component boundaries make it straightforward to run one step in isolation, which is the property that determines iteration speed
  • Serialisable pipeline definitions turn configuration into something you can review, diff and version
  • Deliberately production-oriented, with less abstraction churn than the fastest-moving alternatives

Cons

  • A smaller ecosystem than LangChain, so an unusual source system may be a connector you write yourself
  • The explicit component contract is more upfront ceremony than a simple retrieve-and-generate pipeline warrants
  • Still an orchestration layer, so a complex flow puts your control flow inside it, with less momentum in agent-shaped workloads than the frameworks chasing that market

Best for: Teams that want an explicit, serialisable, testable pipeline definition and value stability over having the newest abstractions.

Pricing: Open source with no licence cost; deepset offers a commercial platform and support separately.

DSPy

DSPy homepage

DSPy is not an orchestration framework and should not be compared to one. You declare modules with typed input and output signatures instead of writing prompt strings, then an optimiser searches over prompts and few-shot examples against a metric you define. Its most useful property is coercive: you cannot get value out of it without an evaluation set, which is the discipline most RAG projects are missing.

Pros

  • Replaces hand-tuned prompt strings with declared signatures and automated optimisation against your own metric
  • You cannot use it without a metric and a labelled set, which forces the evaluation work teams otherwise defer indefinitely
  • The compiled result is portable across models in a way a hand-tuned prompt is not, so a model swap becomes a recompile rather than a rewrite
  • A small conceptual surface compared with the large frameworks, and it composes with whatever retrieval you already have

Cons

  • Requires a labelled evaluation set before it does anything for you, and optimisation runs consume inference budget scaling with the search
  • A research-shaped programming model that is a real mental shift for a team used to editing prompts directly
  • It optimises the part of the pipeline that is usually not your bottleneck, and does nothing for parsing or chunking

Best for: Teams that already have a labelled evaluation set and want prompt and few-shot selection optimised against it instead of hand-tuned by whoever last touched the file.

Pricing: Open source with no licence cost; the real cost is the inference consumed by optimisation runs.

RAGFlow

RAGFlow homepage

RAGFlow is an end-to-end open-source system rather than a library, and it is built around document understanding: layout-aware parsing, chunking templates per document type, and a UI for inspecting how a document was actually broken up. That last feature is the review step from the parsing section, shipped as a product, which is unusual and valuable.

Pros

  • Treats document parsing and layout understanding as the core of the product rather than a preprocessing afterthought
  • Chunking templates per document type acknowledge that a contract, a manual and a spreadsheet should not be chunked identically
  • A UI for inspecting parsing and chunking results, which is the inspection step nearly every team skips
  • Self-hostable end to end, so documents never leave your infrastructure

Cons

  • An opinionated end-to-end system, so integrating it into an existing application is coarser-grained than importing a parser
  • Self-hosting means running its full stack including the datastores underneath, not a single container
  • Less flexible when your retrieval requirements diverge from its pipeline model, and a younger project with thinner accumulated production experience

Best for: Teams that want document understanding, chunking and retrieval as one self-hosted system, with a UI for verifying what the parser actually did.

Pricing: Open source with no licence cost plus a managed offering; self-hosting costs the infrastructure the full stack requires.

Vectara

Vectara homepage

Vectara is retrieval as a managed service: ingest documents through an API and get grounded, cited answers back, with embedding, indexing, reranking and generation handled inside the product. It is the option that removes the entire pipeline this article describes from your responsibility, and the tradeoff is exactly what you would expect.

Pros

  • One API from ingestion through grounded generation, so a small team gets a working pipeline without choosing an embedding model, an index or a reranker
  • Citation and grounding are part of the product rather than something you assemble, and cited answers are what most enterprise buyers actually require
  • No vector database, embedding pipeline or reranker to operate, scale or tune
  • Retrieval quality becomes the vendor’s problem, which is the right division of labour when your differentiator is elsewhere

Cons

  • A managed service, so documents and queries leave your infrastructure and a residency requirement can end the evaluation
  • The pipeline is opaque in the places you might need to tune — chunking strategy, embedding model, reranker choice
  • Usage metering on the retrieval path means cost scales with product usage, and migrating away means building from scratch the pipeline you never had to build

Best for: Teams that need grounded, cited answers over their own documents quickly and have no appetite for owning a retrieval stack.

Pricing: Usage-based on documents stored and queries served, with enterprise tiers above.

Unstructured

Unstructured homepage

Unstructured is an ingestion and parsing platform, not an orchestration framework, and it is worth being blunt about that because it appears on every RAG framework list as if it were interchangeable with LangChain. It extracts documents across a wide range of file types into structured elements — titles, narrative text, tables, list items — and connects to the systems those documents actually live in. What you do with the output is entirely yours.

Pros

  • Focuses on the step that determines RAG quality most, across a genuinely wide range of file types
  • Emits structured elements rather than a flat string, which is the precondition for structure-aware chunking downstream
  • Connectors for the source systems documents live in, which is otherwise a long tail of unglamorous integration work
  • Composes with any framework or with none, because its output is data rather than a control flow

Cons

  • It is not an orchestration framework, so retrieval, reranking and generation remain entirely your design
  • Document processing at volume is compute-heavy, and the hosted API means your documents transit a third party
  • Extraction quality varies by document type, and the boundary between the open-source library and the commercial platform is worth checking against your file types before committing

Best for: Any team whose corpus is heterogeneous file types and whose retrieval is limited by extraction quality rather than by orchestration.

Pricing: Open-source library at no licence cost, plus a commercial API and platform metered on documents or pages processed.

Semantic Kernel

Semantic Kernel homepage

Semantic Kernel is Microsoft’s SDK for orchestrating model calls, plugins and function invocation, with first-class .NET support alongside Python and Java. In a category that treats .NET as an afterthought, that alone decides it for a lot of enterprises. Its model is plugins and functions rather than chains, which maps onto existing service interfaces instead of asking you to restructure code.

Pros

  • First-class .NET support, which almost nothing else in this category offers seriously
  • The plugin and function-calling model maps onto existing service interfaces rather than demanding a new architecture
  • Fits the dependency-injection and configuration patterns enterprise .NET applications already use
  • Integrates with the surrounding identity, governance and AI services a Microsoft-committed organisation already runs

Cons

  • Retrieval features are thinner than the RAG-focused frameworks, and parsing and chunking — where quality comes from — are not its strength
  • The abstractions have changed substantially across versions, so long-lived code needs ongoing maintenance attention
  • Gravity toward one cloud’s services means the smoothest documented paths assume that platform, and there are fewer worked RAG examples than in the Python-centred ecosystem

Best for: .NET enterprises on Azure who want an SDK that fits their existing service and dependency-injection patterns, with retrieval assembled around it.

Pricing: Open source with no licence cost; the spend is the model inference and cloud services it orchestrates.

How to choose

The order here matters more than the shortlist, because steps one to three usually change which framework you need.

1. Read your parsed text. Twenty representative documents, printed and read by a human. If tables have become prose or reading order is scrambled, stop — nothing downstream will fix it, and your framework decision is premature.

2. Build a labelled retrieval set and measure recall at k. Fifty to a hundred real questions paired with the chunk that should answer each. This one measurement tells you whether you have a retrieval problem or a generation problem, and it is the difference between fixing something and redecorating.

3. Add hybrid retrieval and a reranker before you add a framework. Both are bounded changes with measurable effects, and they frequently close the gap you were going to solve with orchestration.

4. Take libraries and connectors freely; adopt an orchestration layer reluctantly. Only when the flow genuinely branches, needs durable resumption, or coordinates multiple agents. Otherwise write the fifty lines.

5. Run the debuggability test before adopting anything that owns control flow. Force a bad answer and try to inspect the retrieval query, candidate scores, reranked list and the exact prompt sent. If you cannot, that framework will cost you an incident later.

6. Change one thing at a time and re-measure. Two simultaneous changes with one metric tells you nothing about either.

ProductWhat it actually isOwns your control flowPicks itself when
LlamaIndexParsing and OCR products plus a library and workflow layerOnly if you adopt the workflow layerParsing quality is your bottleneck
LangChainConnector breadth plus a graph runtime and a commercial tracing platformYes, if you use the chains or graphYou need many integrations and genuinely branching flow
HaystackExplicit component-graph pipeline frameworkYes, explicitlyYou want a serialisable, testable pipeline definition
DSPyPrompt and few-shot optimiser against a metricNoYou already have a labelled evaluation set
RAGFlowEnd-to-end self-hosted document understanding and retrieval systemYes, it is the systemYou want parsing, chunking and retrieval as one deployable product
VectaraManaged retrieval and grounded generation serviceNo, it replaces the pipelineYou need cited answers fast and will not own a stack
UnstructuredIngestion and parsing platformNoHeterogeneous file types and extraction is the limit
Semantic KernelOrchestration SDK, .NET-firstYesYou are a .NET enterprise on Azure

Frequently asked questions

Do I need a RAG framework at all?

For a single retrieve, rerank, generate pipeline, no. That is a parser, an index client, a reranker call and a prompt template — code you can read in one screen and debug without documentation. Take libraries for the parts that are genuinely tedious, especially parsing and connectors. Reach for an orchestration layer when your control flow branches, needs to survive a crash, or coordinates several agents.

Why is my model producing wrong numbers that are present in the source document?

Almost always parsing. Print the extracted text for that document and look at what happened to the table: the numbers survive, the column headers and row boundaries do not, and the model is reading a scrambled sequence of digits. This is not a prompt problem and no amount of instruction fixes it.

What chunk size should I use?

There is no portable answer, and the question is usually the wrong one. Small chunks retrieve precisely and lose context; large chunks carry context and dilute the embedding. Rather than tuning the number, chunk on document structure and use parent-document retrieval — embed the small chunk, return the enclosing section — which sidesteps most of the trade.

Is fine-tuning an alternative to RAG?

They solve different problems. Fine-tuning changes behaviour, format and style; retrieval supplies facts the model does not have and that change without warning. If your complaint is that answers are wrong about your data, retrieval is the fix. If your complaint is that answers are correct but in the wrong shape or register, fine-tuning is a reasonable answer.

How do I tell whether retrieval or generation caused a bad answer?

Measure chunk-level recall against a labelled set. If the correct chunk was not in the retrieved set, it is a retrieval problem and no prompt change will help. If it was retrieved and the answer still contradicts it, it is a generation or context-assembly problem — check whether the chunk survived truncation and where it landed in the prompt.