Buyer’s Guide

Best LLM Evaluation Tools

Written by Govind Kumar Lohar. Reviewed for technical accuracy by Deepak Gupta and Bhaskar Suthar on · Review panel

  • llm
  • evaluation
  • llmops
  • agents

Independent buyer’s guide. No vendor paid to be included, ranked or described a particular way. Written for engineers, architects and the people who sign off on their tooling budget. Editorial policy.

Almost every team shipping an LLM feature has the same evaluation setup, and almost nobody describes it honestly. It is a script that runs twelve examples, asks a frontier model whether each answer looks good, prints an average, and gets run manually when somebody remembers. The average is 0.87. Nobody knows whether 0.87 is good, whether it was 0.87 last month, or what would have to happen for it to be 0.6.

The reason this persists is that LLM evaluation has an unusual property: the measuring instrument is the same kind of unreliable component as the thing being measured. A judge model has biases, drifts when the provider updates it, and prefers text that looks like its own. You are using a ruler made of rubber, which is fine — rubber rulers are still useful — but only if you have measured the rubber.

So this article spends most of its length on the mechanics rather than the vendors, because the mechanics are where the value is. Offline evaluation and online scoring answer different questions and most teams need both. Judge reliability is a solved problem in method and an unsolved one in practice, because the fix requires human labels and nobody wants to produce them. Regression gating in CI works only with a fixed dataset, and the most common failure is a suite that passes because it changed. Red teaming is a different activity from quality measurement and using one for the other leaves you exposed.

Key takeaways

  • Offline evaluation on a fixed dataset gates a change. Online scoring of production traffic catches the drift your dataset does not contain. They are not substitutes.
  • LLM-as-judge has position bias, verbosity bias, self-preference and provider-side drift. The only real fix is a human-labelled calibration set you re-run the judge against, so you measure the judge before trusting its scores.
  • A CI gate needs a fixed dataset, a versioned scoring function and a threshold that fails the build. A suite that changes alongside the code passes for the wrong reason.
  • Red teaming needs adversarial suites and attack generation, not golden datasets. A high quality score says nothing about whether your agent can be talked into calling the refund tool.

Offline evaluation and online scoring answer different questions

Two activities, frequently sold as one feature, with genuinely different requirements.

Offline evaluation runs a fixed dataset through a candidate configuration and scores the results. The dataset does not change between runs, which is what makes the comparison meaningful: version B scored higher than version A on the same inputs, so the difference is attributable to the change. It answers “is this change safe to ship”, and it is the only mechanism that answers that, because production traffic has not seen the change yet. It requires a stable set of inputs with expected outputs, scoring functions, and version control over both — and it can tell you nothing about what your dataset does not contain, which is most of what your users will do next month.

Online scoring runs scorers against live production traffic: sampled or complete, synchronous or in a background job, cheap heuristics or a judge model. It answers “is quality changing”, and it is the only mechanism that answers that, because your dataset was assembled from what you knew at the time. Its constraint is that there is usually no ground truth — you have the input and the output and nothing saying what the right answer was — so online scorers are mostly reference-free: did the answer stay faithful to the retrieved context, did it answer the question, did it emit valid JSON, did the tool sequence terminate, did the user immediately rephrase and ask again.

You need both because of a predictable failure. Offline suites are assembled from known failure modes, so they get better at catching problems you already had and stay permanently blind to new ones — your dataset contains last quarter’s incidents. Meanwhile the traffic mix shifts, a new customer segment arrives with different phrasing, a provider updates the model under a version alias, a retrieval corpus grows and dilutes, and quality moves for a reason your gate cannot see. Online scoring detects the shift; offline evaluation is where the fix gets verified.

The loop closes when they feed each other. An online scorer flags a low-scoring trace, a human confirms it is genuinely bad, and that trace becomes a case in the offline dataset — so the offline suite now covers the failure and will catch its recurrence. That promotion path from production trace to dataset row is the single most important workflow feature in any of the tools below, and it depends on your LLM traces actually containing the payload.

One practical note on cost and latency. Online scoring with a judge model doubles your model calls for every scored request, and the judge’s latency lands on your request if you run it synchronously. Sample, run it asynchronously off the trace, or use cheap deterministic scorers online and reserve the judge for offline runs and a small sampled slice. Judge traffic is a real line on the invoice and is routinely uncounted, which is one of the reconciliation gaps in LLM cost tracking.

LLM-as-judge is a measuring instrument you have not calibrated

Everyone uses a judge model. Very few teams have any evidence their judge agrees with a human, which means their quality metric is an unvalidated proxy being used to make shipping decisions. Here is what goes wrong, specifically, and what the fix actually is.

The biases are systematic, not random

Position bias. In pairwise comparison, judges prefer one position over the other regardless of content. Present A then B and A wins more than it should; swap them and the preference partly follows the position rather than the text. Random noise would average out. This does not — it is a consistent tilt, so a naive pairwise eval can report that your new prompt is better when what it detected is that your new prompt was placed second. The mitigation is cheap and mandatory: run every comparison both ways and count only agreements, treating disagreements as ties. It doubles your judge cost and it is not optional.

Verbosity bias. Longer answers score higher, roughly independent of whether the extra length adds anything. A prompt change that makes the model more expansive will show up as a quality improvement. This is the most common false positive in the whole practice, and the reason you should always chart output length alongside score — if score and length moved together, you have learned nothing about quality and you have learned something about your token bill.

Self-preference. A judge prefers text produced by itself or by models from its own family, with recognisable stylistic markers. If you are choosing between two providers’ models and judging with one of them, the result is compromised in a direction you can predict. Use a judge from a third family, or judge with both and require agreement.

Sensitivity to the rubric. Judges are bad at fine-grained scales. “Score this 1 to 10” produces clustered, poorly discriminating results and shifts with prompt wording. Binary or three-point decisions against explicit criteria are far more stable: does the answer contradict the context, yes or no. Decompose a vague quality question into several sharp binary ones and aggregate, rather than asking one model for a nuanced number.

Order and content of few-shot examples. The examples in your judge prompt shape its scoring, so a judge prompt is itself a versioned artifact that needs the same change control as any other prompt — the argument for treating it as one in prompt management.

Judge drift is a supply-chain problem

Your judge is a model call to a provider. If you pinned a version alias rather than an immutable identifier, the model behind it can change and your entire historical score series becomes non-comparable — silently, with no event in your system to correlate against. A quality “regression” appears, engineering spends three days looking for a change that does not exist, and the cause was upstream.

Three defences, all mundane. Pin the most specific model identifier the provider offers. Re-run a frozen reference set against the judge on a schedule and alert on movement, so a judge change shows up as its own signal instead of contaminating everything else. And record the judge model, its version and the judge prompt version on every score you store, so a series is filterable by which instrument produced it.

The only real fix is a human-labelled calibration set

This is the part teams skip, and skipping it is what makes the rest theatre.

Take 100 to 200 production cases spanning the range of quality you actually see, including the ambiguous middle rather than only obvious passes and failures. Have humans label them against a written rubric — the same rubric your judge gets. Use at least two labellers on an overlapping subset and compute their agreement with each other first: if humans agree only 60% of the time, your rubric is underspecified and no judge can do better than your definition allows. Fix the rubric before you go further. This step alone resolves more evaluation confusion than any tool purchase.

Then run your judge against that labelled set and report agreement with the human labels, per class. Not overall accuracy, which flatters you when the classes are imbalanced — where does the judge say “fine” when a human said “bad”, and how often. That false-negative rate is the number that determines whether your judge can be a release gate, because a judge that misses a third of real failures is a gate with a third of the bar missing.

Then re-run the judge against the calibration set on a schedule, and again whenever the judge model, judge prompt or rubric changes. The calibration set is your standard weight — how you know the rubber ruler has not stretched.

Two closing points. A judge is another model call, so it fails the way model calls fail: rate limits, timeouts, refusals to score sensitive content, malformed output where you expected a label. Handle that explicitly as “unscored” rather than letting it become a zero, because zeros from infrastructure failures look exactly like quality collapse. And prefer a deterministic check wherever one exists — schema validation, required field presence, label-set membership, exact match on extraction, string containment for citations. Deterministic scorers are free, instant, unambiguous and never drift, and a large fraction of what teams use judges for is properly a schema check.

Needs first-hand data: Build a 150-case human-labelled calibration set for one feature, with two labellers and a written rubric. Report inter-labeller agreement first, then per-class agreement between your judge and the human consensus, then repeat the whole measurement with the judge model swapped for one from a different provider family. Publish the false-negative rates. That table is what tells you whether your quality number can gate a release, and essentially nobody has it.

Regression gating in CI

A gate that works has exactly four parts, and the failure mode for each is specific.

A fixed dataset, versioned in or alongside your repository. Fixed is the load-bearing word. The most common broken setup regenerates or extends its test set as part of the run, so a change to the code changes the questions being asked, and the suite passes because it changed rather than because the system is fine. Version the dataset, review changes to it in pull requests like any other test fixture, and require a stated reason to add or remove a case. Growth is good; silent mutation is not.

A scoring function that is code and is versioned. Whether it is a schema check, a string comparison or a judge invocation, it lives in your repository, it has a version, and changing it is a reviewed change. A scorer that changes silently produces a score series that is not a series.

A threshold that fails the build. If the result is a report nobody blocks on, you have a dashboard, not a gate. Two thresholds beat one: an absolute floor on aggregate score, plus a per-case regression check that fails if any previously passing case now fails. The second catches more real problems, because averages absorb individual breakage.

A runtime and cost the team will tolerate. A suite that takes forty minutes and costs real money on every pull request gets disabled within a month, and disabled suites are worse than absent ones because everyone believes they are protected. Structure it in tiers: deterministic checks on every commit; a small judge-based smoke set on pull requests; the full suite nightly and before release. Cache aggressively — if the prompt, model, parameters and input are unchanged, do not re-run the case.

Two things make CI evaluation harder than ordinary testing.

Non-determinism. Even at temperature zero, output is not guaranteed identical across runs, so a single-run pass or fail on a marginal case produces a flaky gate, and flaky gates get ignored. Use thresholds with enough margin that ordinary variance does not flip them, or run marginal cases several times and score the majority. Track your flake rate, because a suite whose failures are 30% noise is one the team has correctly learned to ignore.

Cases go stale invisibly. Expected outputs written against last year’s model may now be wrong, or trivially easy, and a suite of trivially easy cases passes forever while telling you nothing. Review periodically for cases that have never failed and reference answers that are now questionable. Also pin the model version in CI, or your build can fail because a provider shipped an update overnight and you will spend the morning bisecting commits for a change that is not there.

Needs first-hand data: Run your CI evaluation suite ten times against an unchanged commit and record how many cases flip pass-to-fail across runs. That flake rate determines whether the threshold can be tight enough to catch a real regression, and it is the first thing I would measure before letting an eval suite block a merge.

Red teaming is a different activity

Quality evaluation asks whether your system does its job well. Red teaming asks whether it can be made to do something else. These need different data, different metrics and often different tools, and conflating them leaves a real gap.

The material difference is the dataset. Quality evaluation uses a golden dataset: representative inputs with known good outputs, drawn from what users actually do. Red teaming uses adversarial suites, generated rather than collected, and deliberately unrepresentative — the point is to cover an attack surface, not a traffic distribution. Direct prompt injection in the user turn. Indirect injection through content your RAG pipeline retrieves, which is the vector most teams have not thought about at all, because the malicious instruction arrives inside a document the model was told to trust. Attempts to extract the system prompt. Jailbreaks via role-play, hypotheticals, encoding and language switching. Multi-turn escalation, where each turn is individually innocuous. And for agents, the one that actually matters: coercing a tool call with attacker-chosen arguments, which turns a text-generation problem into an unauthorised-action problem.

The metrics differ too. Quality is an average over a distribution. Red teaming is a coverage-and-worst-case exercise, so mean attack success rate is close to meaningless and one reproducible way to make your agent issue a refund is a finding regardless of the other 999 attempts. The output of a red-team run is a list of successful attacks with reproductions, not a score. It is also continuous rather than a one-time audit: your attack surface changes every time you add a tool, expand a retrieval corpus, change a model or loosen a system prompt, so the suite belongs in CI alongside the quality suite.

One boundary worth drawing clearly: red teaming is testing, guardrails are runtime enforcement. Finding an injection that works tells you nothing about whether anything would have blocked it in production, and the enforcement half is covered in LLM guardrails tools. A team with excellent red-team coverage and no runtime controls has a very good description of its own vulnerabilities.

Braintrust

Braintrust homepage

Braintrust is evaluation-first and the most complete answer here for teams whose problem is gating changes. Datasets, scorers, experiments and production logs are one object graph, so the path from a bad production trace to a dataset row to a scored comparison is a supported workflow rather than glue code. Scorers are code you write and version, which matters because the useful metrics for your product are not on anyone’s default menu.

Pros

  • Per-case experiment diffs surface the change that improved the average while breaking three cases you care about, which is the failure aggregates hide
  • Scorers are versioned code rather than a fixed library, so deterministic checks and judge calls live side by side under review
  • Production log to dataset row is a first-class action, which is the loop that keeps an offline suite from going stale

Cons

  • Commercial and managed only, so evaluation data and the payloads it contains live with the vendor
  • Requires the team to have opinions about scoring before it pays off — the tool will not tell you what good looks like
  • Red teaming is not the product; you will need a separate adversarial suite

Best for: Teams that want every prompt, model or pipeline change gated on a scored per-case comparison against a versioned dataset.

Pricing: Usage-based metering on logged spans and evaluation runs with a seat component, and enterprise agreements above that.

LangSmith

LangSmith homepage

LangSmith is LangChain’s commercial evaluation and observability platform — LangChain and LangGraph are open source, LangSmith is not. Its strength for evaluation is the annotation and dataset machinery: annotation queues route traces to human reviewers, labels feed datasets, and datasets drive both offline experiments and online evaluators. That human-in-the-loop path is exactly what a judge calibration set needs.

Pros

  • Annotation queues make human labelling a real workflow, the prerequisite for calibrating a judge and the step most teams never operationalise
  • Offline experiments and online evaluators share dataset and scorer definitions, so the two halves stay consistent
  • Deep LangGraph integration means agent evaluation can target individual nodes rather than only end-to-end output

Cons

  • Not open source; self-managed deployment is an enterprise arrangement, so evaluation data containing customer payloads sits with the vendor
  • The strongest features assume LangChain or LangGraph; on a bespoke stack it is competent but undifferentiated
  • Concentrates framework, runtime, prompts, tracing and evaluation in one vendor, and red teaming is absent

Best for: Teams on LangGraph who need human annotation feeding both offline gates and online evaluators without building the workflow themselves.

Pricing: Per-seat subscription plus usage-based metering on traces ingested, with an enterprise tier for self-managed deployment.

Langfuse

Langfuse homepage

Langfuse is the open-source option covering tracing, prompts, datasets and evaluation in one self-hostable product. For evaluation it gives you dataset runs, custom and model-based scores attached to traces, and human annotation — with the significant advantage that all of it can live inside your own infrastructure, which matters because evaluation datasets are built from real customer payloads and are as sensitive as the traces they came from.

Pros

  • Self-hostable, so evaluation datasets built from real user input never leave your perimeter — the constraint that eliminates most alternatives in regulated environments
  • Scores attach to traces, so online scoring and offline dataset runs share one data model and query surface
  • Prompt version, trace and score sit together, so a quality change is attributable to a specific prompt version

Cons

  • Experiment comparison is functional rather than refined; eval-first tools show per-case regressions more clearly
  • No red-teaming capability, so adversarial testing needs a separate tool
  • Judge calibration is something you build on top of it rather than a guided workflow

Best for: Teams that need evaluation data to stay in their own infrastructure and want it on the same platform as tracing and prompt versioning.

Pricing: Open source with no licence cost when self-hosted, plus managed cloud metered on ingested events with retention tiers.

Arize Phoenix

Arize Phoenix homepage

Phoenix is Arize’s open-source tracing and evaluation project, and its evaluation library is unusually well suited to RAG: reference-free evaluators for whether an answer is grounded in the retrieved context, whether retrieved documents were relevant, and whether a citation is supported. Because retrieval spans carry document identifiers and scores, a failure separates cleanly into “retrieval brought the wrong thing” versus “generation ignored the right thing”. Arize has announced a new chapter with Dynatrace, and Phoenix continues to ship as Arize Phoenix.

Pros

  • RAG-specific evaluators — groundedness, context relevance, citation support — address the failure mode most production LLM apps actually have
  • Runs locally with almost no setup, so evaluation is available while you are developing rather than only after you deploy
  • Open source and free to self-host, so evaluation of sensitive payloads stays local

Cons

  • The commercial roadmap now sits inside an observability incumbent following the new chapter with Dynatrace, so an independent-trajectory assumption no longer holds
  • Phoenix is the open-source slice; production monitoring, drift detection and enterprise controls belong to the commercial Arize platform
  • CI gating is something you assemble around it rather than a first-class feature with thresholds and build integration

Best for: Teams debugging and evaluating RAG pipelines who want groundedness and retrieval-quality evaluators running locally or self-hosted.

Pricing: Open source with no licence cost for Phoenix, with the commercial Arize platform sold on enterprise agreements rather than a public self-serve meter.

Confident AI

Confident AI homepage

DeepEval is the open-source evaluation framework and Confident AI is the platform around it, and the framework’s design choice is the interesting part: evaluations are written as unit tests, executed by a test runner, with assertions and thresholds. That maps evaluation onto machinery your team already understands, which is a large practical advantage — a suite that looks like tests gets run and maintained like tests.

Pros

  • Test-runner ergonomics make CI integration trivial and the mental model familiar, which is the difference between a suite that survives and one that is abandoned
  • A wide library of prebuilt metrics covering RAG groundedness, relevance, hallucination and task completion, so you are not writing scorers from zero
  • The open-source framework runs entirely locally, and it includes adversarial and red-teaming capability

Cons

  • Prebuilt metrics are judge-based, so they inherit every bias above and need calibrating against human labels before you trust them as gates
  • The framework and hosted platform are separate decisions, and the division of features between them can be confusing when planning adoption
  • Broad metric coverage encourages reporting many numbers rather than defending a few, producing dashboards nobody acts on

Best for: Engineering teams that want evaluation as pytest-style tests in CI, running locally, with a broad metric library to start from.

Pricing: Open source with no licence cost for the DeepEval framework; the Confident AI platform adds a usage- and seat-based managed tier.

Ragas

Ragas homepage

Ragas is a focused open-source library for evaluating retrieval-augmented pipelines, and its value is conceptual as much as practical: it decomposes RAG quality into separately measurable components — is the answer faithful to the retrieved context, is it relevant to the question, was the retrieved context relevant and sufficient — so a single quality number becomes a diagnosis. Several metrics are reference-free, which makes them usable on production traffic where no ground truth exists.

Pros

  • Decomposes RAG quality into faithfulness, answer relevance and context quality, turning one unhelpful score into an actionable one
  • Reference-free metrics work on live traffic, so the same library serves offline runs and online scoring
  • A library rather than a platform, so it drops into whatever harness you have with no account or vendor, and it can generate a synthetic test set from your own corpus

Cons

  • A library and nothing more: no UI, dataset management, experiment history or CI reporting — you build all of that
  • Metrics are judge-based and therefore carry judge cost, latency and bias; the numbers are not free and are not calibrated for you
  • Scoped to RAG, so agent trajectories and multi-turn behaviour are out of scope, and metric definitions have shifted across versions so pin the version if you are trending

Best for: Teams with an existing evaluation harness who want well-defined RAG metrics as a library rather than adopting a platform.

Pricing: Open source with no licence cost; the cost is the judge model calls each metric makes and the harness you build around it.

Promptfoo

Promptfoo homepage

Promptfoo is a configuration-driven CLI: declare providers, prompts, test cases and assertions in a YAML file, run it, get a comparison matrix. That makes it the fastest path from nothing to a working gate in CI, and it is also the strongest red-teaming tool here, with adversarial attack generation across injection, jailbreak and extraction categories. Promptfoo is now part of OpenAI, and it still ships the open-source red-teaming and evaluation CLI.

Pros

  • Declarative YAML plus a CLI means a working CI gate in an afternoon, with no SDK adoption and no platform account
  • Assertions span deterministic checks, similarity and judge-based grading, so cheap checks carry most of the load
  • Red teaming is first-class with generated adversarial suites, covering the activity most quality tools ignore, and everything runs locally

Cons

  • Now part of OpenAI, which is a genuine procurement consideration if you are evaluating competing providers with a tool owned by one of them
  • Config-file-driven evaluation gets unwieldy at scale; large matrices in YAML are hard to review and harder to refactor
  • No managed dataset curation, annotation queues or long-term experiment history, so the production-trace-to-dataset loop is manual

Best for: Teams that want a local, declarative evaluation and red-teaming gate in CI without adopting a platform.

Pricing: Open source with no licence cost for the CLI, with a commercial enterprise offering for team features and larger red-teaming deployments.

Patronus AI

Patronus AI homepage

Patronus positions around automated evaluation with purpose-trained evaluator models rather than general-purpose judges, plus hallucination detection and adversarial testing. The interesting architectural claim is that a smaller model trained specifically to detect unsupported claims can be cheaper and more consistent than prompting a frontier model to do it — which, if it holds for your domain, changes the cost calculation for online scoring at volume.

Pros

  • Purpose-trained evaluators are designed to be cheaper and more consistent per score than prompting a frontier model, which matters when scoring production traffic continuously
  • Hallucination and groundedness detection are the design centre rather than one metric among thirty
  • Managed evaluators mean you are not maintaining judge prompts, a real ongoing cost in a hand-rolled setup

Cons

  • A managed evaluator is a black box you cannot inspect or tune to your domain, so calibrating it against human labels is more important here, not less
  • Sending payloads to a third party for scoring is an additional data path to justify, on top of your observability vendor
  • Vendor-defined metric semantics mean your quality definition is partly outsourced, which is uncomfortable when it becomes a release gate

Best for: Teams scoring production traffic at volume who want managed hallucination and groundedness evaluators rather than maintaining judge prompts.

Pricing: Usage-based metering on evaluation calls with tiered plans, and enterprise agreements for higher volume.

Galileo

Galileo homepage

Galileo is now part of Cisco, and that is the first thing to say about it, because it changes what you are buying. The product covers evaluation, guardrail-style runtime protection and agent observability with its own metric suite, aimed at enterprises putting agents into production. As part of Cisco it arrives with enterprise procurement attached, which is an advantage if that is already your paper and a slower path if it is not.

Pros

  • Covers evaluation, runtime protection and agent observability in one platform, reducing the number of vendors touching your payloads
  • Agent-specific metrics target trajectory and tool-selection correctness rather than only final-output quality, which is where agent failures live
  • Can be bought under an existing Cisco relationship rather than as a new vendor, which for some organisations is the deciding factor

Cons

  • Now part of Cisco, so the independent-startup framing is gone and roadmap direction follows a large incumbent’s priorities
  • Enterprise posture means a sales cycle before serious evaluation, slow if you wanted to test something this week
  • Proprietary metric definitions make your quality numbers vendor-shaped and harder to reproduce if you leave

Best for: Enterprises deploying agents that want evaluation, runtime protection and agent monitoring under one contract, particularly where Cisco is already a supplier.

Pricing: Enterprise agreements with usage-based components rather than a public self-serve list price.

LangWatch

LangWatch homepage

LangWatch combines tracing, evaluation and an optimisation layer, with an open-source core you can self-host. For evaluation its distinguishing feature is scenario-based simulation of agent behaviour: rather than scoring single responses, it exercises multi-turn interactions against defined scenarios, which is the only way to catch failures that emerge over a conversation rather than in one turn.

Pros

  • Scenario-based multi-turn simulation tests behaviour that single-response evaluation structurally cannot reach
  • Open-source core with self-hosting, so evaluation data built from customer conversations can stay in your infrastructure
  • Evaluators run both offline and against live traffic from the same definitions

Cons

  • Smaller community than the leaders, which matters most when self-hosting and you are the first to hit a problem
  • Breadth means depth varies, and the optimisation layer assumes more ML literacy than most product teams have
  • CI gating is assembled rather than a polished workflow, and default evaluators need domain adaptation and human calibration before they can gate anything

Best for: Teams building multi-turn agents who need scenario-level behavioural evaluation and want to self-host it.

Pricing: Open source and self-hostable at infrastructure cost, with a managed tier metered on traced volume and evaluation usage.

How to choose

Start from three artifacts you either have or do not.

A fixed dataset. Twenty cases is enough to start and far better than nothing. Draw them from real traffic, include every failure that has caused an incident, and version it in your repository.

A written rubric with human labels. One hundred to two hundred cases, two labellers, inter-labeller agreement computed before you look at any judge. This is what converts your quality number from a guess into a measurement, and it is the artifact teams skip.

A deterministic scorer for whatever can be checked deterministically. Schema validity, field presence, label-set membership, citation containment. Usually more of your quality bar than people expect, and free, instant and immune to drift.

Then route by situation. If your bottleneck is shipping confidence, buy the eval-first platform: Braintrust, or LangSmith on LangGraph. If you need a gate in CI this week and no platform, use Promptfoo — with the caveat that it is now part of OpenAI, which matters if you are choosing between providers. If evaluation data cannot leave your infrastructure, you are choosing among Langfuse, Phoenix, DeepEval, Ragas and self-hosted LangWatch, and that constraint decides more evaluations than any feature. If your system is RAG, start with Ragas or Phoenix for the decomposed metrics. If it is a multi-turn agent, single-response scoring misses most failures, so LangWatch or Galileo. If your concern is security rather than quality, that is red teaming — Promptfoo or DeepEval, plus runtime enforcement from the guardrails tools, because testing and blocking are different jobs.

ToolOffline gateOnline scoringRed teamingSelf-hostable
BraintrustCore strengthYesNoNo
LangSmithYes, with annotationYesNoEnterprise only
LangfuseYesYesNoYes
Arize PhoenixYes, RAG-focusedYesNoYes
Confident AI / DeepEvalTest-runner nativeVia platformYesFramework runs locally
RagasLibrary onlyReference-free metricsNoYes, it is a library
PromptfooDeclarative, CLI-nativeLimitedCore strengthYes, runs locally
Patronus AIYesManaged evaluatorsYesNo
GalileoYesYes, with runtime protectionYesEnterprise deployment
LangWatchYes, scenario-basedYesLimitedYes

One judgment to close on. I would rather have twenty hand-labelled cases, a schema check and a threshold that fails the build than a platform reporting fourteen judge-based metrics nobody has calibrated. The tooling here is good and getting better. The bottleneck is almost never the tool.

Frequently asked questions

How many test cases do I need to start?

Twenty is enough to catch obvious regressions and vastly better than none. A hundred starts to give stable aggregate scores. What matters more than count is coverage of the failures you have actually seen — every incident should leave a case behind, so the suite grows toward your real risk rather than an arbitrary target.

Can I trust LLM-as-judge scores?

Only relative to a human-labelled calibration set you have measured the judge against. Judges have position bias, verbosity bias and self-preference, and they drift when the provider updates the model. Run pairwise comparisons in both orders, pin the most specific model identifier, judge with a model from a different family than the one you are evaluating, and re-check agreement with human labels on a schedule.

What is the difference between evaluation and guardrails?

Evaluation measures quality, offline or on sampled production traffic, and informs decisions. Guardrails enforce policy at runtime, in the request path, and block or modify individual requests. An evaluation suite tells you an injection attack works; only a guardrail stops it.

Why did my evaluation scores change when I did not change anything?

Most likely your judge model changed under a floating version alias, your dataset was regenerated as part of the run, or you are seeing ordinary non-determinism on marginal cases. Rule them out in that order: pin the model identifier, confirm the dataset is fixed and versioned, then run the suite ten times against an unchanged commit and measure the flake rate.