Buyer’s Guide

Best LLM Cost Tracking and FinOps Tools

Written by Govind Kumar Lohar. Reviewed for technical accuracy by Deepak Gupta and Bhaskar Suthar on · Review panel

  • llm
  • cost
  • llmops
  • infrastructure

Independent buyer’s guide. No vendor paid to be included, ranked or described a particular way. Written for engineers, architects and the people who sign off on their tooling budget. Editorial policy.

The question that starts every one of these projects is the same, and it sounds trivial: which customers are costing us money on AI?

Then somebody opens the provider dashboard and discovers it has one number. A monthly total, maybe split by model, maybe split by API key if you were disciplined about keys. Nothing about customers, nothing about features, nothing about which agent step burned four thousand tokens re-reading a document it had already read twice. The finance team wants a per-account margin. What exists is a single figure and a graph of it going up.

This is not a reporting gap you can fill with a better tool. It is a data-collection gap, and the data was never collected. A provider bills you per token against an API key; it has no idea that this particular call was made on behalf of your enterprise customer, from your summarisation feature, on step three of a retry loop. If that context was not attached at the moment the request was made, no dashboard, no spreadsheet and no vendor can reconstruct it afterwards.

I have watched teams spend a quarter trying to work backwards from invoices and timestamps. It does not work. The fix is upstream and it is cheap, but only if you do it before you need the answer.

Key takeaways

  • Attribution is a tagging problem at the call site. Metadata not attached when the request is made is gone forever, and every tool below is downstream of that decision.
  • A cache changes your unit economics before it changes your invoice, and you cannot manage either without hit rate measured per feature.
  • A hard budget limit means a request gets refused, which means your product needs a defined behaviour for “budget exhausted”. Teams that skip designing that never turn enforcement on.
  • LLM-native tools can attribute per customer because they sit at the call. Cloud FinOps platforms read the invoice, so they can allocate an account and never a customer.

Attribution is a tagging problem you solve at the call site

This is the paragraph that matters more than the rest of the article combined, so I will be blunt about it. Every unit of cost analysis you will ever want must be attached to the request as metadata at the moment you make it. Not derived later, not joined from another table, not inferred from timestamps. Attached, at the call site, before the request leaves your process.

The minimum useful set, and I would treat this as a checklist for code review on any file that calls a model:

  • Customer or tenant identifier. The one thing finance will ask for and the one thing nobody adds first. Without it there is no per-account margin, ever.
  • Feature or product surface. “Summarisation”, “support-reply-draft”, “onboarding-extract”. This is what lets you kill an expensive feature that nobody uses, which is usually the single largest saving available.
  • Environment. Staging and CI traffic hitting production keys is a recurring surprise, and undifferentiated environment tags mean you spend a week arguing about whether the spike was real.
  • Agent step or node name. In an agent, the aggregate cost is meaningless; you need to know that the planner is cheap and the tool-selection retry loop is not.
  • User identifier or session. For per-seat products, and for finding the one power user distorting your averages.
  • Prompt version. So a cost increase that coincides with a prompt change is attributable rather than mysterious. This is the same link that makes quality regressions traceable, which is why prompt management and cost tracking end up in the same conversation.
  • Request purpose. Whether this call was a first attempt, a retry, a judge evaluating another call’s output, or a background job. Judge calls in particular are invisible in most accounting and can be a third of your spend.

Three practical notes.

Put it in one place — a wrapper function, a gateway or a client subclass that every model call routes through. Tags applied at forty individual call sites will be inconsistent within a month, and inconsistency is worse than absence because it produces confidently wrong reports. If no such chokepoint exists, creating one is the actual first task, and it is the prerequisite for guardrails, caching and routing later.

Use stable, low-cardinality values. Feature names come from an enum, not a string literal typed by whoever wrote the handler. Customer identifiers are necessarily high cardinality and that is fine; “summarize”, “summarization” and “summarisation” as three features is self-inflicted.

Propagate through the agent. A tool call made on behalf of a request must carry the originating call’s customer and feature tags. This is the same context-propagation problem that breaks LLM traces into fragments, with the same cause: an async task started without the ambient context.

Do this and the tool choice becomes easy, because you have the data and are only choosing where to aggregate it. Skip it and every tool here shows you the total you already had.

Cached and uncached calls are two different products

Caching is the most effective cost lever available to an LLM application and the one most likely to be reported wrongly.

Two mechanisms, often confused. Provider-side prompt caching is where the provider stores the processed prefix of your prompt — typically a long system message, a tool schema, or retrieved documents — and charges a reduced rate for those input tokens on subsequent calls that share the prefix. You still make the call, you still get a fresh completion, and your invoice still has a line for it. Application-side caching, semantic or exact-match, is where you never make the call at all: a cache in front of the provider returns a stored response, so the invoice line does not exist.

The economics differ sharply. Prompt caching reduces the input cost of a call whose output you still pay for in full, which helps most when your prompts are long and your completions short — RAG with big contexts, agents with large tool schemas. Response caching removes the cost entirely and introduces a correctness question, because you have decided that a similar-enough question deserves an identical answer.

Here is the part that trips teams up. A cache changes your unit economics immediately and your invoice not at all until you can see the hit rate per feature. If overall spend is flat while traffic doubled, your cache is working and your reporting is hiding it. If hit rate is 60% in aggregate, that number is nearly useless — it might be 95% on a docs-lookup feature and 4% on the conversational surface where each turn is genuinely unique, and the average tells you to invest in the wrong place.

So the minimum instrumentation is: cache hit or miss, and which kind, tagged per feature; cached versus uncached input token counts recorded separately, because they price differently; and cost-avoided computed as what the call would have cost, which is the only number that justifies the cache’s existence to anyone. Any tool that reports a single blended cost per request has thrown away the distinction you need. Semantic caching has its own failure modes and its own trade space, covered in the semantic caching guide.

Needs first-hand data: For one week, record cache hit rate and cost-avoided separately for each product surface, alongside a human spot-check of fifty cache hits on the conversational surface to count how many returned an answer that was similar enough to retrieve but wrong enough to matter. The hit-rate number is worthless without that second figure, and nobody publishes both.

Alerting tells you afterwards; enforcement refuses the request

There is a hard line between the two and most teams sit on the wrong side of it without noticing.

Alerting watches spend and notifies you when it crosses a threshold. It is easy to enable, it is honest about what it does, and it is a detection control — by the time the alert fires, the money is spent. For a monthly budget with a slow-moving product this is adequate. For a feature that a single script can drive into five figures overnight, it is not a control at all.

Enforcement means the tool refuses the request. Some component — a gateway, a proxy, your own wrapper — checks the remaining budget for a key, a customer or a virtual key, decides it is exhausted, and returns an error instead of calling the provider. This is a preventive control and it is the only thing that actually caps exposure.

And this is where teams stall, for a reason that has nothing to do with tooling: enforcement means your product must have a defined behaviour for “budget exhausted”, and nobody has designed one. The gateway is ready. The application is not, because the code path assumes a model response always arrives. Turn enforcement on without that design and your first budget breach is a 500 in a user’s face.

The behaviours worth designing, roughly in order of how often they are the right answer:

Degrade to a cheaper model. The request succeeds on a smaller model and quality drops. Best default for most consumer-facing features, and it requires that your prompts and output parsing actually work on the fallback — verify that, because a prompt tuned on a frontier model frequently produces unparseable output on a smaller one. That verification is an evaluation task against a fixed set, which is why evaluation needs wiring in before you rely on this.

Queue and defer. The request is accepted, the user is told it is processing, and it runs when budget resets or a human approves an override. Correct for batch work, wrong for anything interactive.

Refuse with a real message. A specific error — the feature is at its limit for this period, here is what to do — surfaced in the UI rather than a generic failure. Unglamorous and often correct, particularly for internal tools and a free tier where the limit is the point.

Serve stale from cache. Return the last good response, marked as possibly out of date. Only viable where staleness is tolerable, and it makes your cache a reliability component.

Two design points people miss. Budget checks must be atomic under concurrency, or a hundred simultaneous requests each see remaining budget and all proceed — the same double-spend problem as any counter, and the reason enforcement lives in a gateway with a shared store rather than per-process state. And a request’s cost is unknowable until after the response, because output tokens are not known in advance, so enforcement always reserves an estimate then reconciles.

Enforcement, routing and fallback are gateway responsibilities, which is why the AI gateways guide and this article overlap so much — for most teams the cost control and the gateway are the same component.

Your token accounting will not match the provider’s

It never does, and the gap is information rather than noise. Reconcile monthly, per model, and treat a persistent unexplained delta as a bug in your instrumentation.

The usual causes, in roughly the order they explain the difference:

Retries you did not count. An SDK with automatic retries on a timeout or a rate-limit response may have made three attempts. If the first two produced partial generation before failing, they may still be billable, and your application logged one call. Client-side retry configuration is a cost setting, and almost nobody treats it as one.

Streaming responses aborted mid-flight. A user closes the tab, your timeout fires, your load balancer cuts the connection at thirty seconds. Generation continued or partially completed on the provider’s side and you are billed for what was produced, while your code recorded a failure with no token count. Long-running agent calls behind an aggressive proxy timeout make this a systematic error, not an occasional one.

Cached-input pricing tiers. If some input tokens billed at a cache-read rate and your accounting applies one input price to all of them, your estimate is high. If you are on the wrong side of a cache-write surcharge, it is low. Either way you must record cached and uncached input tokens as separate fields, or reconciliation is impossible in principle.

Embedding calls nobody counted. Ingestion pipelines, re-indexing jobs and vector-store backfills make enormous numbers of cheap calls. They rarely go through the same wrapper as your chat path, so they are absent from your accounting and very present on the invoice. Re-embedding a corpus after a chunking change is a large one-off nobody attributes to the feature that motivated it.

Judge and evaluation calls. Online scoring, LLM-as-judge, guardrail classifiers and reranking are all model calls. Running through a different code path with different tags, they land in the invoice as unexplained volume. Tag them explicitly as evaluation traffic, because they are a real and often surprising fraction of spend.

Agent loops that retry quietly. An agent that fails to parse a tool result and tries again has doubled the cost of that step with no error surfacing. Loop-iteration count per request is a cost metric, not only a quality one.

Tokenisation and time boundaries. A client-side tokeniser will not exactly match the provider’s count, and invoices close on their calendar boundary in their timezone. Neither explains a 20% gap; both explain a 2% one, and knowing that keeps you from chasing the wrong thing.

Needs first-hand data: Reconcile one month of your own token accounting against the provider invoice, per model, and attribute every percentage point of the difference to a named cause from the list above. Record what fraction came from retries, aborted streams, uncounted embeddings and judge calls. That breakdown is the most useful artifact in this whole exercise and it is specific to your architecture.

LLM-native tools and cloud FinOps platforms are not competing

Before the product blocks, the distinction that decides which half of this list you should even read.

LLM-native tools sit at the call. A gateway, proxy or SDK wrapper sees every request, so it can read your metadata, count tokens on both sides, know whether the response was cached, and attribute cost to a customer, a feature and an agent step. They can also enforce budgets, because they are in a position to refuse. What they generally cannot do is tell you about your GPU instances, your vector database bill, your object storage or the rest of the cloud spend the feature also incurs.

Cloud FinOps platforms read the bill. They ingest cloud provider billing exports and increasingly AI provider spend, then allocate it across accounts, tags, teams and products, and put it next to compute, storage and network. They give you total cost of ownership and a shared vocabulary with finance. What they structurally cannot do is attribute an invoice line to your end customer, because that information was never in the invoice. A FinOps platform can tell you the AI line item grew 40% and which cloud account it sits in; it cannot tell you it was one enterprise trial account running a document pipeline.

Most teams past a certain size need both, and the sequencing is: LLM-native first, because per-customer attribution is the question people are actually asking, and it is the one that expires if you do not instrument for it. The pattern is the same cost-shock story that drives teams off general observability platforms, one layer down — the mechanics in the Datadog alternatives piece are worth reading precisely because the shape repeats: a meter you did not model, growing on an axis you did not expect, with attribution that arrives too late to prevent it.

Helicone

Helicone homepage

Helicone is a proxy: change your provider base URL and every call is logged with cost, tokens, latency and whatever custom properties you attach as headers. That header-based tagging is exactly the call-site attribution this article is about, and it works without an SDK. It is LLM-native, open source, and it can run inside your own infrastructure. Helicone has announced it is joining Mintlify.

Pros

  • Custom properties as headers give per-customer and per-feature attribution with a one-line change and no SDK adoption
  • Being in the request path means caching, rate limits and per-key spend limits are available rather than a separate integration
  • Open source and self-hostable, keeping prompt payloads and cost data inside your perimeter

Cons

  • A proxy in the hot path is an availability dependency; the async logging mode avoids that and gives up the enforcement and caching that made it attractive
  • Tagging is only as consistent as the discipline applied to it, and a proxy cannot invent a customer identifier nobody sent
  • Joining Mintlify moves the roadmap into a documentation company, so stop reading it as an independent observability startup

Best for: Teams that want per-customer and per-feature LLM cost attribution across several services with a base-URL change rather than an instrumentation project.

Pricing: Open source and self-hostable at infrastructure cost, with a managed tier metered on logged requests and retention, and cached traffic metered separately from uncached.

Portkey

Portkey homepage

Portkey is now Prisma AIRS AI Gateway, part of Palo Alto Networks, and generally available for enterprises. As a cost tool it is the most complete gateway-side answer here: virtual keys with their own budgets and rate limits, per-request metadata, routing and fallback, caching, and spend visible per key, model and tag. The consequence of the new ownership is worth stating plainly — this is enterprise security-vendor software now, which helps if your procurement already runs through that vendor and changes the conversation if you were shopping for an independent startup.

Pros

  • Virtual keys with individual budgets turn enforcement into configuration rather than application code, which is the only way most teams ever ship a hard limit
  • Metadata travels with the request into cost reporting, so per-customer attribution is the default rather than an add-on
  • Routing, fallback and caching sit in the same component as the budget, so “degrade to a cheaper model when the budget is exhausted” is expressible as policy

Cons

  • The independent-product framing is gone; this is enterprise software from Palo Alto Networks, with the sales cycle that implies
  • A gateway in the request path is another hop to operate and becomes a single point of failure for every AI feature you have
  • Pricing power now sits with a large incumbent, which is a renewal consideration rather than a technical one

Best for: Enterprises that want per-key budget enforcement, attribution and governance in one gateway and are comfortable buying it from a security incumbent.

Pricing: Enterprise commercial gateway with usage-based metering on requests routed and tiered enterprise agreements rather than a public per-request list price.

LiteLLM

LiteLLM homepage

LiteLLM is the open-source route to the same capability: an SDK and a proxy server that normalise many providers behind one interface, with per-key virtual budgets, spend tracking, tags and rate limits built into the proxy. It is what I would reach for when the requirement is “we need budget enforcement and per-team attribution and we are not buying anything this quarter”.

Pros

  • Virtual keys with budgets, spend tracking and rate limits are in the open-source proxy, so hard enforcement needs no commercial contract
  • Normalises many providers behind one call shape, so attribution metadata has one place to live even across providers
  • The SDK-only mode gives cost accounting without a network hop in the request path, which is a genuinely useful middle option

Cons

  • You operate it: the proxy plus its state store lands on your critical path, and enforcement under concurrency depends on that store being healthy
  • Provider abstraction is leaky at the edges, and newly released provider parameters lag the provider’s own SDK
  • Reporting is functional rather than a finance-grade surface, so you will export the data elsewhere to present it

Best for: Engineering teams that want per-key budget enforcement and multi-provider cost attribution from software they run themselves, with no vendor in the payload path.

Pricing: Open source with no licence cost for the SDK and proxy, plus a commercial enterprise tier for support and administrative features; self-hosting converts the cost to infrastructure and operator time.

Langfuse

Langfuse homepage

Langfuse is an observability platform first, and cost is derived from the traces it already collects — token counts and model per span, priced up and rolled into per-trace, per-user, per-session and per-tag totals. That derivation is the point: because cost sits on the same object as the prompt, the completion and the score, you can ask which expensive traces were also the bad ones, which no billing tool can answer.

Pros

  • Cost lives on the same trace as quality scores and payloads, so “expensive and wrong” is a filterable set rather than two investigations
  • Agent-step attribution comes free from the span tree, which is where aggregate per-request cost stops being useful
  • Open source and self-hostable, keeping cost and payload data inside your infrastructure

Cons

  • Observability, not enforcement: it can tell you a budget was exceeded and cannot refuse a request, so hard limits still need a gateway
  • Cost is computed from token counts and a price table, so it is an estimate to reconcile against the invoice rather than a ledger
  • Calls that produce no trace — background embedding jobs, anything outside the instrumented path — are absent from the totals

Best for: Teams that already want LLM tracing and evaluation, and want cost attributed to the same traces rather than in a separate financial tool.

Pricing: Open source with no licence cost when self-hosted, plus managed cloud metered on ingested trace volume with retention tiers.

OpenMeter

OpenMeter homepage

OpenMeter is now OpenMeter by Kong, and it solves an adjacent problem that gets conflated with this one: metering usage so you can bill your own customers for it. It ingests usage events, aggregates them against entitlements and plans, and exposes balances you can enforce against. If your requirement is “charge customers for AI usage and cut them off at their plan limit”, this is the category, not a cost dashboard.

Pros

  • Purpose-built for usage-based billing, so token events become customer-facing meters, entitlements and balances rather than internal charts
  • Entitlement balances are a real enforcement primitive at the customer level, which is the missing half of most budget conversations
  • Being part of Kong means it sits alongside a gateway that can enforce what the meter says rather than only reporting it

Cons

  • Not an LLM cost tool: it does not know model prices, does not compute provider spend, and will not tell you your margin unless you feed it both sides
  • Adopting it is a billing project with product, pricing and finance stakeholders, a much larger commitment than a cost dashboard
  • Requires call sites to emit clean, deduplicated usage events, so it sits downstream of the same instrumentation work

Best for: Products that resell AI usage and need customer-facing metering, plan entitlements and hard limits, especially teams already on Kong.

Pricing: Open source core with a commercial managed offering metered on ingested usage events, sold under the Kong commercial model.

Vantage

Vantage homepage

Vantage is a cloud FinOps platform that has extended into AI provider spend, and it belongs here for what no LLM-native tool does: putting model spend next to compute, storage, data transfer and vector database costs, in the cost-allocation vocabulary your finance team already uses. It reads billing data, which is both its strength and its hard limit.

Pros

  • Total cost of ownership in one view — GPU instances, managed vector stores, egress and model spend — which is the number that decides whether a feature is profitable
  • Cost allocation, showback and chargeback across teams and accounts is mature, because that is the core product rather than a new module
  • Provider-agnostic across cloud accounts, so a multi-cloud AI footprint consolidates

Cons

  • Reads invoices and billing exports, so it can allocate to an account or tag and structurally cannot attribute to your end customer, feature or agent step
  • No enforcement; it will not refuse a request, so it is detection only
  • Cannot distinguish cached from uncached calls or a retry from a first attempt, because the invoice does not carry that

Best for: Platform and finance teams that need AI spend allocated alongside all other cloud costs in one FinOps surface with existing chargeback workflows.

Pricing: Subscription scaled to the volume of cloud spend under management rather than to your LLM call volume.

CloudZero

CloudZero homepage

CloudZero is the other side of the same category, with a sharper focus on unit economics — cost per customer, per product, per feature — assembled from billing data plus allocation rules and your own telemetry. That focus makes it the closest a FinOps platform gets to the question this article opens with, and the caveat is that it gets there by ingesting the attribution data you produce.

Pros

  • Built around unit-cost questions — per customer, per feature, per environment — rather than account-level totals
  • Allocation of shared and untaggable infrastructure is handled explicitly, which is where naive tag-based reporting gives up
  • Combines cloud infrastructure and AI provider spend, so a feature’s model cost and serving cost sit in one margin calculation

Cons

  • Per-customer AI attribution still depends on you emitting per-customer usage from your call sites, so it consumes the instrumentation work rather than removing it
  • Detection only, with no ability to enforce a budget or refuse a call
  • Blind to LLM-specific structure — cache hits, retries, judge calls, loop depth — because none of it appears in billing data

Best for: Organisations that need defensible cost-per-customer and cost-per-feature economics across cloud and AI spend, with finance and engineering working from one model.

Pricing: Enterprise subscription scaled to cloud spend under management, with no public self-serve list price.

OpenRouter

OpenRouter homepage

OpenRouter is a hosted routing layer across many models and providers behind one API, and it earns a place here because it changes the shape of the cost problem rather than reporting on it. Spend consolidates into one account with per-key attribution and per-model pricing visible in one place, and switching a feature to a cheaper model is a parameter change rather than an integration.

Pros

  • One account and credential across many models, so per-key and per-app spend attribution works without you running infrastructure
  • Model prices and availability are comparable in one surface, making “is this feature worth a frontier model” a cheap experiment
  • Fallback across providers is built in, so a provider outage needs no routing logic of your own

Cons

  • A commercial intermediary now sits between you and the provider, so availability, rate limits and data handling are partly its concern
  • Attribution granularity is bounded by what you can express per key and per request; it does not replace your own customer and feature tagging
  • Consolidating spend into one vendor account also consolidates a dependency, and enterprise procurement will ask about it

Best for: Small and mid-size teams that want multi-model access, per-key spend attribution and cheap model substitution without running a gateway.

Pricing: Pay-as-you-go against prepaid credit at per-model token rates with a margin on provider pricing, rather than a platform subscription.

How to choose

Do the instrumentation first. Everything below assumes the tags exist.

Before you evaluate anything: find or create the single chokepoint every model call routes through, and attach customer, feature, environment, agent step, prompt version and request purpose. Include embedding pipelines and judge calls, which is the step everyone skips. This work is not wasted under any tool choice, and no tool substitutes for it.

Then pick by the question you are actually being asked.

If it is “which customers are unprofitable”, you need an LLM-native tool reading your call-site metadata — Helicone, LiteLLM, Portkey or Langfuse. A FinOps platform cannot answer it from invoices.

If it is “what does this AI feature cost us in total”, you need a FinOps platform, because model spend is often the smaller half once GPU serving, vector storage and egress are counted.

If it is “how do we stop this happening again”, you need enforcement, which means a gateway — LiteLLM if you will run it, Portkey if you will buy it — and first a product decision about what happens when the budget is gone.

If it is “how do we charge customers for this”, you are in metering and billing, not cost tracking.

ToolCategoryAttributes per customerCan refuse a request
HeliconeLLM-native proxyYes, from request headersYes, via key limits in proxy mode
PortkeyLLM-native gatewayYes, from request metadataYes, per virtual key budget
LiteLLMLLM-native SDK and proxyYes, from tags and virtual keysYes, per virtual key budget
LangfuseLLM observabilityYes, from trace metadataNo — reporting only
OpenMeterUsage metering and billingYes, for your customers’ usageYes, via entitlement balances
VantageCloud FinOpsNo — allocates accounts and tagsNo
CloudZeroCloud FinOps, unit economicsOnly from data you supplyNo
OpenRouterHosted model routerPer key, not per customer nativelyYes, via credit exhaustion

Two closing judgments. I would not buy a FinOps platform to answer a per-customer LLM question; the categories are not substitutes and the disappointment is expensive. And I would not enable hard enforcement until the degraded path has been evaluated against a fixed test set — a cheaper model producing unparseable output is a worse outage than the overspend you were preventing.

Frequently asked questions

Why can’t I just work backwards from the provider invoice?

Because the invoice holds a total against an API key and nothing about your customers, features or agent steps. There is no join key, and timestamps do not help because concurrent traffic from many customers interleaves. Attribution must be attached at the call, and if it was not, that period is permanently unattributable.

Does a semantic cache reduce my bill?

Only where it prevents a call. An application-side cache returning a stored response removes the invoice line entirely. Provider-side prompt caching reduces the price of input tokens on a call you still make and still pay output for. You need cached and uncached token counts as separate fields to see either clearly.

Where should budget enforcement live?

In a gateway or proxy with a shared state store, not in application code. Budget checks must be atomic under concurrency or simultaneous requests will each see available budget and collectively blow through it. One shared component also means every service inherits the policy instead of each team implementing it differently.

How much should my token accounting differ from the invoice?

A small single-digit percentage is normal, from tokenisation differences and billing-period boundaries. Anything larger has a findable cause: retries you did not log, streams aborted after generation started, embedding pipelines outside your instrumented path, or judge and guardrail calls tagged as something else.