Every agent framework demo is a happy path. Five steps, tidy summary, applause. The question that separates these products is what happens on step seven of a ten-step run — the one where step four already charged a customer’s card, step six returned malformed JSON, and the pod running the loop got evicted mid-inference.
Most frameworks answer that by not answering it. The loop is a while in one process, the state is a list of messages on the heap, and a restart takes all of it with you. Fine for a chat turn: the user sees an error, hits retry, nothing lost but seconds. Unacceptable for a workflow that calls a payment API, because the retry is not free — the side effects already happened and the framework does not know which ones.
So read this category on three questions: where does the state live, can you see what the model was actually sent, and can a run stop for a day and start again. That is why a workflow engine older than the current wave of LLM tooling belongs in an article about agent frameworks. Choose on durability first; the abstraction is the cheap part to change later.
Key takeaways
- The real axis is state. An in-memory loop loses everything on restart — fine for a chat turn, wrong for anything with side effects.
- Resumable and restartable are different guarantees. Restarting only works if every step is idempotent, and tool calls that move money never are.
- If you cannot print the prompt as actually sent and the verbatim arguments of every tool call, you cannot debug the agent at 3am.
- Approval gates that pause for hours are a persistence problem. Frameworks that model them as in-process callbacks break on the next deploy.
- MCP turns tools into a deployable surface with an owner. It does not make any tool safe to expose, and every server widens what prompt injection can reach.
Step seven: what your framework does when the process dies
First, name the state. A run holds more than a conversation: message history, tool results, a plan or scratchpad, retrieved documents, token accounting, and the position in the loop. A framework that persists “the messages” has persisted a transcript, not a run.
There are three genuinely different answers, and the labels get used loosely.
In-memory. The loop runs in your process. Some frameworks add conversation persistence so a follow-up turn has history — a different feature, surviving the turn rather than the crash. If your agent is a chat endpoint with read-only tools, stop here; more durability costs you a state store to operate.
Checkpointed. After each step the framework serialises run state to a store under a thread identifier, and a resume loads the last checkpoint and continues. Two details decide whether it is real: granularity, and whether the checkpoint lands before or after the side effect. Written after the model decides to call a tool but before the tool runs, it will re-run that tool on resume. Written after the tool returns, it will not — but a crash between the call and the write leaves you unable to tell which happened.
Durable execution. Workflow code is deterministic and every non-deterministic thing — an LLM call, an HTTP request, a clock read — goes through an activity the engine records in an event history. Recovery replays your code against that history, fast-forwarding past work already done. Strongest guarantee available, and not an LLM idea at all. The cost is a programming model with rules: no wall clock in workflow code, no unrecorded I/O, explicit versioning when you change a workflow with runs in flight.
The distinction that matters operationally is resumable versus restartable. Restartable means running the whole thing again and hoping — fine when every step is idempotent, an incident when step four is POST /charges. The fix is not framework-specific: every write tool takes an idempotency key derived from run identifier plus step index, enforced downstream. Do that and replay is a cost problem instead of a correctness problem. Skip it and no checkpointer saves you, because the framework cannot know two identical POSTs mean two charges.
Replaying an LLM call is also neither free nor deterministic: a different model version can produce a different plan. Durable engines record the response as fact rather than re-asking, which means the event history contains prompts and completions — a retention decision to make on purpose.
Needs first-hand data: Build the same seven-step workflow with one side-effecting tool on each shortlisted framework, then
SIGKILLthe process between step four and step five. Record whether the run resumes, from which step, whether the side-effecting tool executes twice, and how much state (plan, retrieved documents, token counts) survives versus only the message list.
You cannot debug what you cannot see
You need the prompt as actually sent — not your template, but the final payload after tool schemas were serialised, after the framework’s own hidden preamble, and after history pruning or summarisation ran. Nearly every genuinely confusing agent bug lives in that gap: the framework trimmed the message carrying the instruction, or rendered an enum as prose, or injected a formatting demand that fought yours.
You need tool calls verbatim: the raw argument string before parsing, the parsed object, the result, and the error when parsing failed. “Tool call failed” without the arguments has thrown away the only evidence.
You need per-step accounting, because agent cost is a per-run number nobody forecast. A ten-step run that resends the whole transcript each step is quadratic in tokens, and that is the default in more frameworks than you would guess.
And you need replay from a step: load the checkpoint at step six, change the prompt or a tool result, run forward. Without it, iterating on step six means paying for steps one through five every time.
This is why agent tooling converged on spans rather than log lines: a run is a tree, and a span tree is the same shape. Where those spans go is the subject of the LLM observability roundup; the framework’s job is producing them at all.
Approval gates are a persistence problem
A gate that pauses a run until a human approves has to survive the wait, and the wait spans a deploy, a scale-to-zero and probably a Monday.
The wrong shape, which several frameworks ship, is a callback: the loop blocks on a promise that resolves when someone clicks. That works in a notebook and dies the moment the process is not the same process. The right shape is a suspend — the run writes state, exits, and an external event keyed by run identifier resumes it. Same machinery as durable execution, which is why frameworks with a real checkpointer tend to have real approval gates.
Approve the concrete call, not the intent. If the human approves “refund this order” and resumption asks the model what to do next, the model may re-plan and issue a different call. The approved artifact must be the exact tool name and arguments.
Timeouts are a business decision. What happens after 48 hours of silence — fail, escalate, proceed with a safe default? A framework with no timeout primitive pushes that into a cron job you write.
Authorisation is not the gate. Who may approve which action is access control, and it belongs with your other permissions. Same boundary problem as tool least-privilege in the LLM guardrails guide: the framework can pause the run, but it cannot know support may refund up to a limit and a manager beyond it.
Tool definitions and where the boundary should be
A tool is a JSON schema plus a function, and the schema is prompt engineering wearing a type hat. The model picks tools by reading names and parameter descriptions, so a parameter called flag with no description is a reliability bug. Schemas derived from type hints keep one source of truth; hand-written schema dictionaries drift within a sprint.
Then ask what happens when arguments fail validation. Bad behaviour is raising into your application. Good behaviour is returning the validation error to the model as a tool result so it can correct itself, under a retry bound you can set — an unbounded self-correction loop is how a run burns real money producing nothing.
The last question is where the tool runs. An in-process function is versioned with your agent and invisible to every other agent in the company. A tool behind a protocol is a service with an owner and reuse. That trade is what MCP is about.
MCP: the standard underneath, not a product
MCP — the Model Context Protocol — is not one of the frameworks below. It is a wire protocol for exposing tools, resources and prompts to any client that speaks it, with no vendor, no dashboard and no bill. Ranking it against LangGraph is a category error: you pick it underneath a framework, and most frameworks here consume MCP servers as a tool source.
What it changes is ownership. Before, a tool was a function in each agent’s repository, so the fifth team needing “look up a customer” wrote the fifth implementation with the fifth set of bugs. After, it is a server with one deploy and one owner, consumed by agents and by editors nobody on your team wrote — an improvement in the way an internal HTTP API improved on a shared library, and it inherits the same problems.
What it gives you
- Tools become a deployable surface with a version, an owner and a change process instead of duplicated functions
- One server serves every client that speaks the protocol, including runtimes and editors you do not control
- Runtime discovery, so an agent enumerates available tools at startup rather than having them compiled in
- Resources and prompts travel with the tool, so the context needed to use it correctly ships alongside it
What it does not do
- It does not make a tool safe to expose. A
run_sqltool is exactly as dangerous over MCP as it was locally, and now more callers reach it - It carries no authorisation model for your domain. Whose credentials the server uses, and whether the end user behind the agent may use them, is yours to build
- Every server widens the surface prompt injection can reach: text returned by one tool is model input, and model input steers the next tool call. An unfiltered reader of the open web beside a write-capable tool is the classic dangerous pair
- It does not standardise quality. Descriptions, error shapes, pagination and idempotency stay per-server decisions
- Runtime discovery means your effective prompt can change without a deploy on your side
Treat a server like a public endpoint: least privilege on its credentials, a hard allowlist of which servers each agent may reach, and no write tool beside an unfiltered reader unless you accepted that risk deliberately.
LangGraph

LangGraph models an agent as a graph of nodes over an explicit typed state object, and that one decision drives the rest: because state is declared it can be serialised, and because it can be serialised there is a checkpointer interface from in-memory up to Postgres. It comes from the same company as LangChain and LangSmith, so the framework is open source while the commercial observability platform beside it is not.
Pros
- Typed state plus a pluggable checkpointer makes durability a configuration choice rather than a rewrite
- Interrupts are first-class, so approval gates suspend the run and resume from an external event
- Time-travel lets you iterate on step six without re-paying for steps one through five
Cons
- Verbose for simple agents; a three-tool assistant becomes more scaffolding than logic
- Inherits LangChain’s history of abstraction churn, which makes older examples actively misleading
- The polished tracing and evaluation experience is a separate commercial product
Best for: Teams building multi-step workflows with side effects and approval steps who want durability and replay without adopting a general workflow engine.
Pricing: Open source with no licence cost; you pay for model calls, the checkpoint store you run, and separately for the hosted platform beside it.
CrewAI

CrewAI organises work as a crew of role-playing agents — researcher, writer, critic — coordinated sequentially or by a manager agent, with tasks as the unit of work. Its Flows layer adds explicit control when the crew abstraction is too loose. What the metaphor hides is control flow: what looks like delegation between colleagues is a prompt chain, and you debug prompts, not the org chart.
Pros
- Fastest path from idea to working prototype, because roles and tasks match how people describe the problem
- Flows give an explicit event-driven path when crew behaviour is too implicit to debug
- Large prebuilt tool library, so common capabilities do not start from a blank file
Cons
- The role metaphor obscures a prompt-composition pipeline, making failure attribution harder than an explicit graph
- Durability is thinner than a checkpointed model, so long runs with side effects need your own persistence
- Hierarchical delegation multiplies token cost unpredictably, since a manager agent re-reads subordinate output
Best for: Teams prototyping research, content or analysis pipelines that decompose naturally into roles, where no step touches money.
Pricing: Open source core with no licence cost plus a commercial hosted platform on subscription; the dominant cost is tokens across several cooperating agents.
AutoGen

AutoGen came out of Microsoft Research on the idea that conversation is orchestration: agents talk, and the pattern of who speaks next is the program. Later versions restructured around an event-driven actor core with a chat layer on top, which made it far more suitable for services than notebooks. It is the strongest option here for open-ended group problem solving and the weakest for anything you need bounded.
Pros
- Conversational orchestration handles iterative problem solving a fixed graph cannot express
- The actor-style core separates message passing from agent logic, a sound base for distributed execution
- Strong support for code-executing agents including sandboxed execution
Cons
- Termination is the hard part of a conversational design: without careful stopping conditions, agents converse until a token budget intervenes
- Heavy API churn across major versions, so many published examples target an architecture that no longer exists
- Durable long runs are not its centre of gravity, so pausing a day for review is work you add
Best for: Research-flavoured and code-generation workloads where the value is agents iterating with each other while a human watches.
Pricing: Open source with no licence cost; spend is model tokens plus whatever you run the sandboxed executors on.
Pydantic AI

Pydantic AI takes the position that agents are ordinary typed Python and the interesting guarantees come from validation. Tool schemas derive from function signatures, outputs validate into models, and a validation failure becomes a message back to the model rather than an exception in your handler. It is the least magical framework here — you can read the control flow without learning a graph DSL.
Pros
- Schemas from type hints mean a refactor cannot silently desynchronise the schema from the function
- Output validation with model-facing retries turns malformed responses into a bounded correction loop instead of a 500
- Integrates with durable execution engines rather than reimplementing them, keeping the framework small
Cons
- Python only, so a TypeScript product team is out
- Multi-agent orchestration is deliberately thin, so complex topologies are yours to compose
- Durability comes from what you wire it into rather than a built-in checkpointer
Best for: Python teams who want typed, testable agents with hard output contracts and would rather compose durability than inherit a graph runtime.
Pricing: Open source with no licence cost; the vendor behind it sells a separate hosted observability product, and your real bill is tokens.
Mastra

Mastra is the strongest TypeScript-native option, which matters more than it sounds: if your product is a Next.js application, an agent framework in the same language removes a service boundary and a deployment story. Its workflow primitive supports suspend and resume, so human approval is modelled rather than improvised.
Pros
- One language and one deployment for your web app and your agents, with types shared across the boundary
- Workflows suspend and resume, so approval gates are a supported shape rather than a callback you keep alive
- Memory, retrieval and evaluation included, which shortens assembly for a first real agent
Cons
- TypeScript only, and the deeper machine-learning tooling ecosystem is still Python
- Younger project with a smaller body of production war stories than the Python incumbents
- Serverless targets and multi-day suspended workflows interact awkwardly; check execution limits first
Best for: Product teams shipping agent features inside an existing TypeScript or Next.js application who do not want a Python service in the path.
Pricing: Open source framework with no licence cost plus an optional commercial cloud for hosting and tracing on a usage-based subscription.
Google ADK

The Agent Development Kit is Google’s code-first agent framework, designed to run locally and deploy onto the Gemini Enterprise Agent Platform — the product formerly called Vertex AI. Framework and managed runtime were designed together: sessions, state, artifacts and evaluation have a local implementation for development and a managed one for production, from the same code.
Pros
- The same code runs locally and managed, with session and state services swapped by configuration
- Multi-agent composition, sessions and artifact handling are built in rather than assembled
- Ships an evaluation harness and native MCP support, so neither is a separate adoption decision
Cons
- Strong gravity toward Google’s platform: model-agnostic in principle, best supported on Gemini
- Durability and scaling live in the managed runtime, so self-hosting the full experience means rebuilding parts of it
- The platform rename means much older documentation still uses the previous product name
Best for: Teams already on Google Cloud who want agents that deploy onto a managed runtime without writing their own session and state layer.
Pricing: Open source SDK with no licence cost; the managed runtime is usage-based on the platform, metered separately from model consumption.
Strands Agents

Strands Agents is AWS’s open-source agent SDK, and its bet is the opposite of a graph: give the model a prompt, a set of tools and a loop, and let the model drive. There is very little scaffolding, MCP is a first-class tool source, and it targets AWS compute with Bedrock as the smooth model path while staying provider-agnostic.
Pros
- Very small surface area — a model, tools and a loop — which is quick to read and reason about
- Native MCP support rather than an adapter, so external tool servers plug straight in
- Fits AWS serverless and container deployment paths without a bespoke hosting story
Cons
- A model-driven loop trades determinism for flexibility: you cannot guarantee step order, which rules it out for fixed-sequence or regulated processes
- Long-run resumption depends on what you build around it rather than a checkpointer in the framework
- Documentation and examples assume AWS, so the off-AWS path is thinner
Best for: AWS-native teams building tool-using assistants where flexibility matters more than a guaranteed sequence of steps.
Pricing: Open source with no licence cost; you pay AWS for compute and the provider per token for inference.
Temporal

Temporal is not an agent framework and should not be judged as one. It is a durable execution engine: workflows are ordinary code, every side effect goes through a recorded activity, and recovery replays your code against that event history so it resumes exactly where it stopped. It belongs here because the hardest agent problems — surviving a crash at step seven, retrying one activity without redoing the run, pausing three days for approval — are what it was built for long before anyone had an agent loop.
Pros
- The strongest durability guarantee available: runs survive process death, deploys and infrastructure failure with no checkpoint design from you
- Retries and timeouts are per-activity configuration, so a flaky tool retries without re-running the LLM steps around it
- Signals and timers make a gate that waits days a native construct rather than an integration project
- The event history is a real audit trail of every step and side effect
Cons
- You write the agent loop yourself; there are no agent, tool or memory abstractions in the box
- The determinism rules are a genuine learning curve, and violating them produces confusing replay failures
- You operate a cluster or buy the managed service, which is meaningful infrastructure for one agent feature
Best for: Teams whose agents perform real transactions or span days, where a durable audit trail of every side effect is a requirement.
Pricing: Open source server self-hosted at infrastructure cost, plus a managed cloud metered on workflow actions and storage rather than per seat.
OpenAI Agents SDK
The OpenAI Agents SDK is the deliberately minimal option: agents with instructions and tools, handoffs to transfer control, guardrails alongside, sessions that persist conversation, and tracing built in rather than bolted on. It is honest about its scope — a loop with good ergonomics, not a workflow engine. It speaks other providers through the request shape covered in the compatible-proxy guide, so it is less provider-locked than the name suggests.
Pros
- Smallest concept count here — agents, handoffs, guardrails, sessions — so a team is productive on day one
- Tracing is built in and on by default, so the “we cannot see what the agent sent” problem never starts
- Handoffs are a clean routing primitive that avoids a manager-agent prompt
Cons
- Sessions persist conversation, not mid-run step state, so a crash between tool calls loses the run
- Default tracing goes to the vendor’s platform, a data-flow decision to make consciously rather than discover in an audit
- Provider gravity in features and docs even where the abstractions are portable
Best for: Teams shipping a tool-using assistant quickly, where runs are short, effects are reversible, and vendor-hosted tracing is acceptable.
Pricing: Open source SDK with no licence cost; you pay per token for inference, with hosted tracing metered separately on the vendor’s platform.
How to choose
Start from one sentence describing your worst run, not from a framework.
A bad answer is a chat problem. Take the smallest framework your language supports, wire tracing on day one, and spend the time on tool descriptions and evaluation. A graph runtime is overhead; the evaluation tools guide is a better use of next week.
A duplicated side effect makes the framework the second decision. The first is idempotency keys on every write tool, enforced downstream, because that is the part that keeps working when you change frameworks.
A run that waits for a human eliminates every framework whose approval mechanism is an in-process callback. Test it specifically: reach the gate, restart the process, then approve.
A compliance question — what did the agent do to this account eight weeks ago — needs an event history, not traces with 30-day retention. That points at durable execution and at the retention design in the compliance gateway guide.
| Framework | State model | Picks itself when |
|---|---|---|
| LangGraph | Checkpointer, interrupts, time-travel | You need resumption and approval gates without a workflow engine |
| CrewAI | Mostly in-process, Flows for control | A multi-agent prototype this week, no step touching money |
| AutoGen | In-process, conversation-scoped | Open-ended iteration between agents, often over code |
| Pydantic AI | Composed — you bring durability | Readable typed agents with hard output contracts |
| Mastra | Workflow suspend and resume | Your product is TypeScript and you refuse a Python service |
| Google ADK | Session and state services | Production is Google Cloud and you want the runtime included |
| Strands Agents | Bring your own persistence | AWS-native, flexibility beats a fixed step order |
| Temporal | Event-sourced replay of every step | Runs span days, move money, or need an audit trail |
| OpenAI Agents SDK | Conversation sessions only | Runs are short, effects reversible, speed wins |
Needs first-hand data: Run one identical five-tool agent on your two finalists against a fixed set of recorded tasks, and record tokens per completed task, wall-clock latency, and how many runs needed a manual restart. Framework overhead shows up in the token number, not the feature list.
Frequently asked questions
Do I need an agent framework at all?
Often not. A tool loop with three tools and a stopping condition is about eighty lines you fully understand, and easier to debug than any framework. Adopt one when you need something it genuinely provides — durable resumption, approval gates, multi-agent routing, a tracing surface you would otherwise build.
What is the difference between a checkpointer and durable execution?
A checkpointer snapshots run state after steps, so resuming loads the last snapshot and continues. Durable execution records every side effect in an event history and replays your code against it, fast-forwarding past completed work. Checkpointing is simpler and usually enough; durable execution gives stronger guarantees and an audit trail, at the cost of a stricter programming model.
How do I stop a run costing an unbounded amount?
Three hard limits: maximum steps per run, maximum tokens per run, maximum retries per tool — enforced outside the model’s control, because a model asked to stay under budget will not. Routing calls through a gateway puts the enforcement point and the per-run accounting in one place, which is the argument in the AI gateway hub.
Related reading
- Best AI gateways — the layer that gives agent runs shared rate limits, budgets and per-run cost accounting.
- Best LLM observability tools — where agent spans go, and which platforms show the prompt as actually sent.
- Best LLM guardrails tools — tool least-privilege and prompt-injection containment for agents calling MCP servers.
- Best model routing and multi-provider tools — failover for the provider outage that lands mid-run.
- Best RAG frameworks — retrieval as a tool, and why retrieval quality caps agent quality.