Agent frameworks change quickly. As of this review, relevant primary documentation includes LangGraph, CrewAI, Pydantic AI, OpenAI Agents SDK, Claude Agent SDK, Google Agent Development Kit, Hugging Face smolagents, and Microsoft Agent Framework. Re-check maintenance status, supported runtimes, licenses, and release notes when you evaluate them.
This article provides a shortlist and a repeatable evaluation method. Its recommendations are hypotheses to test on your workload, not measured market-share or production-adoption findings.
What an agent framework is for
Before comparing, clarify what we’re choosing among. An “agent framework” typically provides:
- A way to define agents — what’s their role, what tools they have, what’s their behavior.
- An execution loop — call LLM, parse output, decide what to do, call tools, repeat.
- State management — what does the agent remember, how is it organized.
- Tool integration — how tools are defined and exposed.
- Orchestration — multiple agents working together, branching workflows, retries.
- Observability hooks — tracing, logging, debugging.
- Convenience utilities — prompt templates, common patterns, helpers.
Each framework prioritizes these differently. Some are heavy on orchestration; some focus on agent definition; some are minimal layers over the model APIs.
The landscape
LangChain / LangGraph
The big one. LangChain started as a Python library for chaining LLM calls; it became the de facto framework for many AI projects. LangGraph is the agent-specific framework built on top.
What it does well:
- LangGraph for state machines. The graph model (nodes for steps, edges for transitions, state passed between) is well-suited for complex agent workflows.
- Rich ecosystem. Many integrations: vector DBs, model providers, tools, observability.
- LangSmith for observability. Mature tracing and debugging UI.
- Wide adoption. Many examples, much documentation, many people who know it.
What it doesn’t:
- Abstraction tax. LangChain in particular has many layers of abstraction. Debugging is harder; understanding what’s actually happening takes effort.
- API churn. Frequent breaking changes. Code from 12 months ago often needs updates.
- Performance overhead. Layers of indirection cost latency and tokens.
- Learning curve. Real fluency takes weeks.
When to choose:
- Complex agent workflows with branching state.
- Teams that benefit from a standard framework (many engineers, common patterns).
- When you want LangSmith observability.
When to skip:
- Simple chatbots or single-agent loops (direct API is simpler).
- Teams that have been burned by LangChain’s churn previously.
- Projects where every millisecond of latency matters.
CrewAI
A multi-agent framework focused on role-based agents. Each agent has a role, goal, backstory; they collaborate on tasks.
What it does well:
- Multi-agent orchestration. Built-in support for agents talking to each other, delegating, collaborating.
- Role-based mental model. Easy to think about (“the Researcher agent does X; the Writer agent does Y”).
- Simpler than LangGraph for multi-agent. Faster to get started.
- Active community.
What it doesn’t:
- Limited single-agent depth. If your task is a single complex agent, CrewAI’s abstractions can feel mismatched.
- Performance. Multi-agent setups multiply LLM calls; cost and latency scale fast.
- Maturity gap. Younger than LangChain; some rough edges remain.
- Opinionated. Less flexibility than direct frameworks.
When to choose:
- Multi-agent workflows where role differentiation makes sense.
- “Crew” framings (a team of agents working together).
- Prototyping multi-agent ideas quickly.
When to skip:
- Single-agent tasks (overkill).
- Production where performance matters (multi-agent is expensive).
- Tasks where the “agent collaboration” framing is more theater than substance.
Pydantic AI
A newer framework focused on type safety and developer experience.
What it does well:
- Strong typing. Pydantic-based throughout. Inputs and outputs are typed. Errors caught at development time.
- Clean API. Less abstraction than LangChain; closer to the model APIs.
- Modern Python. Async, type hints, Pydantic v2.
- Model-agnostic. Works with most providers.
What it doesn’t:
- Smaller ecosystem. Fewer integrations than LangChain.
- Evidence to obtain. Because the APIs and ecosystem are newer than some alternatives, require pinned upgrade, load, persistence, recovery, and debugging tests rather than asserting a general scale verdict.
- Less orchestration tooling. Not as feature-rich as LangGraph for complex workflows.
When to choose:
- Type-safety-conscious Python teams.
- Single-agent or simple multi-agent setups.
- Teams that prefer minimal abstraction.
When to skip:
- Very complex orchestration (LangGraph might fit better).
- Non-Python projects (it’s Python-only).
- When you need a huge ecosystem of pre-built integrations.
OpenAI Agents SDK
OpenAI’s official agent framework, tuned for OpenAI models.
What it does well:
- Optimized for OpenAI. Tuned for current OpenAI tool-calling and handoff patterns.
- Simple API. Less abstract than LangChain.
- Built-in handoffs. Multi-agent handoffs are first-class.
- Mature tracing. Built-in observability tied to OpenAI dashboard.
What it doesn’t:
- OpenAI lock-in. Designed for OpenAI’s models. Using other providers is awkward.
- Less flexibility. Some patterns are easier in more general frameworks.
- Newer than LangChain. Smaller community.
When to choose:
- All-in on OpenAI models.
- Want a vendor-supported path.
- Simple-to-moderate agent complexity.
When to skip:
- Multi-provider strategy (better to use a more general framework or direct API).
- You’re using Anthropic or Google primarily.
Anthropic Claude SDK
Similar — Anthropic’s path for building agents with Claude.
What it does well:
- Optimized for Claude. Especially good for Claude’s extended thinking, computer use, MCP integration.
- Idiomatic for Claude models.
- Strong MCP support.
What it doesn’t:
- Claude lock-in. Same trade-off as OpenAI Agents SDK.
When to choose:
- All-in on Claude.
- Heavy use of Claude-specific features.
When to skip:
- Multi-provider strategy.
Direct API
Skip frameworks entirely. Call OpenAI / Anthropic / Gemini APIs directly. Write the loop yourself.
What it does well:
- Full control. Every aspect of the system is yours.
- No abstraction tax. What you see is what runs.
- Easy to debug. No layers to dig through.
- Easy to optimize. No framework overhead.
- No version churn. You upgrade when you choose to.
What it doesn’t:
- More code. Patterns the framework handles, you handle.
- Reinventing. Common patterns are reimplemented per project.
- Less standardization. Different teams build similar systems differently.
When to choose:
- Mature teams shipping production systems where reliability matters more than convenience.
- Single-focused use cases that don’t need a framework’s flexibility.
- Performance-critical paths.
- After prototyping with a framework and learning the patterns.
When to skip:
- Greenfield, exploratory, “what should we build” phase (framework helps you discover patterns).
- Teams with limited engineering capacity.
LlamaIndex
Started as a RAG-focused library; has grown into broader agent territory.
What it does well:
- RAG-heavy systems. Best-in-class for retrieval-focused agents.
- Data connectors. Many integrations for data sources.
- Mature retrieval abstractions.
What it doesn’t:
- Agent abstractions are weaker. Better for RAG than general agents.
- Some overlap with LangChain ecosystem.
When to choose:
- Heavy retrieval / RAG focus.
- Need many data source connectors.
When to skip:
- Non-RAG agent work.
Microsoft Agent Framework
Microsoft describes Agent Framework as the successor to both AutoGen and Semantic Kernel and publishes migration guides for each. Evaluate Agent Framework for new work; treat AutoGen and Semantic Kernel as migration inputs rather than the default current recommendation.
What it does well:
- Microsoft ecosystem integration. Works well with Azure, .NET, Microsoft 365.
- Explicit workflows and agents. The documentation separates open-ended agent work from deterministic workflows.
- Migration paths. Microsoft documents migration from both predecessor frameworks.
What it doesn’t:
- Ecosystem fit. Provider and language support must be verified against your required stack.
- Migration state. Teams already using the predecessors need to price the transition and compatibility work.
When to choose:
- Teams aligned with Microsoft’s AI and .NET/Python ecosystem.
- Existing AutoGen or Semantic Kernel users evaluating the documented successor.
The dimensions to think about
Choosing isn’t about picking a winner — it’s about matching tradeoffs to your project.
Dimension 1: Complexity of orchestration
How complex are your agent workflows?
- Simple (chatbot, single agent, linear flow): direct API or Pydantic AI.
- Moderate (single agent, branching logic): Pydantic AI, LangGraph, direct API.
- Complex (multiple agents, state machines, retries): LangGraph, CrewAI, custom.
- Very complex (large state machines, parallel agents, complex routing): LangGraph or custom.
Dimension 2: Production maturity
How important is reliability vs experimentation?
- Experimentation / prototyping: any framework helps you move fast.
- Production, customer-facing: require deterministic control points, trace export, durable state where needed, stable error handling, and a tested upgrade path.
- Production, consequential: choose the smallest dependency surface that passes the workload’s security and reliability acceptance tests. This may be a framework or a direct implementation.
Dimension 3: Team size and skills
- Any team size: score current expertise, on-call ownership, reviewability, and maintenance capacity. Headcount alone does not determine the right abstraction.
Dimension 4: Vendor strategy
- Multi-provider: general frameworks (LangChain, Pydantic AI) or direct API.
- Single vendor: vendor SDKs (OpenAI Agents SDK, Anthropic Claude SDK).
Dimension 5: Performance sensitivity
- Latency- or cost-sensitive: benchmark the same trace with equivalent prompts and tool schemas. Framework overhead may come from extra model calls, serialization, persistence, or tracing; do not assume it exists without measuring.
- Less sensitive: still test error paths and operational behavior, not only the happy-path demo.
Dimension 6: Observability needs
- Strong out-of-box: LangChain + LangSmith.
- DIY: any framework + your own observability layer.
Framework versus direct API
Direct provider calls reduce framework-specific code but do not eliminate orchestration, state, retries, tracing, validation, or upgrade work. A framework can supply some of those primitives; it can also constrain them. Compare both options in the bake-off and include the code your team would otherwise have to own.
Keep model calls, tool contracts, business rules, and persistence behind boundaries you control. That makes a later migration possible without pretending it will be free.
A practical decision framework
If you’re choosing for a new project:
Step 1: Define the project.
- What’s the agent complexity?
- How many engineers?
- Production or prototype?
- Single or multi-provider?
Step 2: Apply heuristics.
| Scenario | Recommended |
|---|---|
| Prototype, complex orchestration | LangGraph |
| Prototype, multi-agent | CrewAI |
| Production, simple agent | Direct API or Pydantic AI |
| Production, complex orchestration | LangGraph or custom |
| Single-vendor (OpenAI / Anthropic) | Vendor SDK |
| Type-safety-focused Python team | Pydantic AI |
| Heavy RAG | LlamaIndex + your choice |
| Microsoft ecosystem or predecessor migration | Microsoft Agent Framework |
Step 3: Prototype, then assess.
Time-box a representative slice for each serious candidate. Use the same fixtures and assess:
- Does it fit your patterns?
- Are you fighting the framework or with it?
- Is debugging tractable?
- Is performance acceptable?
If yes: continue. If no: try another or go direct.
Step 4: Don’t lock in irreversibly.
Even within a framework, structure your code so that swapping is possible. Isolate the framework usage to a thin layer; build your logic in framework-agnostic code.
Patterns that travel across frameworks
Regardless of framework choice, certain patterns are universal:
Separation of concerns. Prompt management separate from agent logic separate from tool definitions separate from execution loop. Each framework helps with some; you handle the rest.
Observability. Trace every LLM call. Trace every tool call. Aggregate metrics. This is your job regardless of framework.
Step budgets and escape hatches. Consequential production agents need externally enforced limits. Verify what the framework supplies and add what is missing.
Eval suites. Frameworks don’t include serious eval tooling. Build them separately (Promptfoo, Braintrust, custom).
Production hardening. Rate limits, idempotency, error handling, fallbacks. Framework gives some primitives; you build the rest.
If you focus on these universal patterns, the specific framework choice matters less. The team’s discipline matters more.

Framework-by-framework verdicts
Use these as shortlist hypotheses, then validate them:
LangChain/LangGraph: graph and durable-workflow concepts may fit explicit state machines. Measure dependency weight, persistence semantics, and upgrade cost.
CrewAI: role-oriented multi-agent abstractions may fit a genuinely collaborative workload. First prove that multiple agents outperform a simpler workflow.
Pydantic AI: a strong candidate for Python teams that value typed model and tool boundaries.
OpenAI Agents SDK / Anthropic Claude SDK: good if you’re committed to that vendor. Otherwise, lock-in risk.
LlamaIndex: a candidate for retrieval-heavy systems; compare its retrieval and data abstractions separately from agent orchestration.
Direct API: a candidate when the workflow is small or the team needs precise control and accepts ownership of the missing orchestration primitives.
Microsoft Agent Framework: the current Microsoft candidate and documented successor to AutoGen and Semantic Kernel.
An illustrative migration sequence
This is a hypothetical scenario, not a reported customer case:
Initial implementation: The team builds a bounded feature with a framework and records its baseline behavior, dependency graph, traces, and upgrade procedure.
Operational hardening: The team adds evaluation, observability, security, persistence, recovery, and performance evidence required by the workload.
Measured friction: A framework abstraction may become a constraint. The team isolates it behind an owned boundary only after a trace, benchmark, or upgrade test demonstrates the problem.
Upgrade decision: The team rehearses a pinned upgrade. If acceptance tests fail or migration cost exceeds the value, it can retain the supported version temporarily, replace a component, or move a measured path to a direct API.
Ongoing ownership: Keep framework components that continue to pass acceptance and security tests; replace only paths where measured benefits exceed migration and maintenance cost.
This is one valid trajectory. Others are valid too — some teams stay in LangChain happily; some skip it from day one.
A different lens: what you’re really choosing
Beyond the framework, you’re choosing:
- A community to learn from.
- A pace of API churn to live with.
- A set of patterns to standardize on.
- A debugging experience.
- An observability story.
- A future migration cost.
The framework is one expression of these. But these are the things that affect your team day-to-day.
A framework that fits your community, your churn tolerance, your patterns, your debugging style, your observability needs — that’s the right choice. Without those fits, even the most popular framework is wrong for you.
The system matters more than the framework
There is no universal best agent framework in 2026. The right choice depends on project complexity, team size, production maturity, vendor strategy, and team preferences.
A working approach:
- Match framework to project requirements using the heuristics above.
- Prototype before committing.
- Structure code so swapping is possible.
- Focus on universal patterns regardless of framework.
- Expect to evolve — what fits now may not fit in 12 months.
Direct API and frameworks are both legitimate production choices. The defensible choice is the one that passes a representative bake-off and has an owned upgrade and exit plan.
Pick the one that fits today. Adjust when it stops fitting. The system you build matters more than the framework you build it with.



