Agent SDKs and Frameworks

Agent SDKs and Frameworks

Definition: An agent SDK — sometimes called an agent development kit, or ADK — is a software library that packages the recurring plumbing of building an LLM-based agent into reusable building blocks: the perceive-think-act loop from Intelligent Agent, tool/schema registration for Function Calling (Tool Use), conversation and session state, structured handoffs between specialized agents from Multi-Agent System, safety guardrails, and observability. A developer using one configures and composes these primitives instead of hand-writing a request/response loop, JSON parsing, retry logic, and error handling from scratch for every new project. An agent SDK sits a layer above a raw model API — it doesn’t replace the underlying completions or messages endpoint a provider exposes, it wraps it in scaffolding that turns “call a model in a loop” into “build a production agent.”

How It Works

What an Agent SDK Actually Provides

  • Loop management — the code that repeatedly calls the model, executes any tool calls it returns, feeds results back, and decides when the task is finished, implementing the ReAct-style reasoning loop described in Function Calling (Tool Use) so a developer never reimplements it by hand.
  • Tool registration — a decorator or schema-builder that turns an ordinary function into a model-callable tool automatically, usually inferring the JSON Schema from the function’s own type signature rather than requiring it to be hand-written.
  • Session and state management — conversation history, working memory across turns, and often hooks into persistent storage so a long-running agent can resume where it left off.
  • Handoffs — a structured, first-class way for one agent to transfer an in-progress conversation to another, more specialized agent, turning the orchestration patterns from Multi-Agent System into a concrete API rather than ad hoc prompt engineering.
  • Guardrails — input and output validation hooks that can halt or redirect an agent before a model call goes out or before its result reaches the user, catching disallowed requests or malformed output at a defined checkpoint.
  • Tracing — built-in logging of every model call, tool call, and handoff, so a run can be replayed and inspected after the fact instead of debugged from application logs never designed for it.
  • Model Context Protocol (MCP) integration — most current-generation agent SDKs ship a built-in MCP client, so an agent built with one can connect to any MCP server and gain its tools without extra integration work.

The Named Players

This is a fast-moving space, and specific feature sets change often — but the broad shape of what’s out there is stable enough to be worth naming:

SDK / FrameworkMaintainerDistinguishing idea
OpenAI Agents SDKOpenAILightweight primitives — agents, handoffs, guardrails, tracing — as explicit first-class concepts, with a built-in MCP client
Google Agent Development Kit (ADK)GoogleOpen-sourced, designed for both a single simple agent and deep multi-agent hierarchies; interoperates with the Agent2Agent (A2A) protocol for cross-vendor agent communication
Claude Agent SDKAnthropicPackages the same underlying harness that powers Claude Code — context management, permissioned tool use, subagents, MCP support — for building general-purpose custom agents, not just coding ones
LangChain / LangGraphLangChainAn earlier, broader ecosystem: LangChain focuses on chaining and composing LLM calls and tools, LangGraph layers an explicit graph-based state machine on top for more deterministic, debuggable multi-step control flow
CrewAICrewAIOpen-source and explicitly role-based — a “crew” of agents is defined with a role, a goal, and a backstory each, oriented toward orchestrator-worker pipelines
AutoGenMicrosoftOpen-source, models agents as conversable entities that exchange messages with each other — one of the earlier influential entrants in this space

“ADK” as a Term, Not Just a Product

Worth being precise about, since the acronym gets used two ways: ADK is both the specific name of Google’s product and, informally, a generic label for “a kit for building agents” — the same way SDK is both a generic software-engineering term (see Software Development Kit (SDK)) and part of many specific product names (the Android SDK, the AWS SDK). A job listing or article that says “experience with an ADK” may mean Google’s product specifically, or may just mean any agent-building framework — the acronym alone doesn’t disambiguate, the surrounding context has to. An agent SDK is, in the general sense from that note, exactly what any SDK is: a packaged toolkit wrapping an underlying capability — here, a model API plus a tool-execution loop — so a developer doesn’t hand-roll it.

Choosing Between Them

  • Vendor-native vs. vendor-agnostic. A provider’s own SDK (OpenAI’s, Anthropic’s, Google’s) is typically the most polished experience for that provider’s own models, at the cost of some lock-in; open frameworks like LangGraph, CrewAI, or AutoGen are more portable across model providers but usually require more manual assembly to reach the same level of polish.
  • How much orchestration the task actually needs. A single, well-scoped agent barely benefits from any of this — see Multi-Agent System’s own guidance on when a single agent remains the better choice — so reaching for a full framework’s handoff machinery on a task that doesn’t need it adds complexity without adding capability.
  • Maturity of tracing and guardrails. For a production system handling real users, the depth of an SDK’s observability and safety tooling often matters more than which specific model it defaults to, since those are the parts hardest to bolt on retroactively.
  • Interoperability needs. If agents built on different vendors’ stacks will eventually need to delegate work to each other, a protocol like A2A (below) matters more than which single SDK any one team picked.

A2A: The Complement to MCP

Model Context Protocol (MCP) standardizes how an agent reaches tools and data. A separate, newer class of protocol — Agent2Agent (A2A), initiated at Google and since opened to broader industry governance — standardizes how one autonomous agent discovers and delegates work to another autonomous agent, potentially built by a different team on a different vendor’s stack entirely. The distinction is real and easy to blur: MCP is agent-to-tool, A2A is agent-to-agent. A travel-booking agent calling a weather API is an MCP-shaped problem; that same agent handing off “handle the payment” to a completely separate payments agent it doesn’t control is an A2A-shaped one.

Why It Matters

  • Standardizes the loop so teams stop reinventing it, incorrectly. Retry logic, context truncation, and tool-call parsing are exactly the kind of plumbing that’s easy to get subtly wrong when every team writes its own from scratch.
  • Makes multi-agent handoffs testable. A handoff implemented as a named SDK primitive can be unit-tested and traced directly; a handoff implemented as an ad hoc prompt instruction usually can’t be, in the same reliable way.
  • Turns “why did the agent do that” into an inspectable log, rather than a guessing game reconstructed after the fact from whatever the application happened to log.
  • Inherits the whole MCP ecosystem for free. An SDK-built agent with a built-in MCP client can connect to any existing MCP server immediately, without the team writing a single integration.
  • Lowers the barrier between a fragile prototype and a genuinely production-grade agent — guardrails, structured retries, and real tracing are the difference between a weekend demo and something a team can safely operate.
  • Cross-vendor delegation is becoming possible without full lock-in. Protocols like A2A mean an agent doesn’t have to be a monolith built entirely on one company’s stack to cooperate with agents that are.
  • Creates a genuine reuse layer. A well-built custom tool, guardrail, or handoff pattern, once wrapped in the SDK’s abstractions, can be reused across many different agents the same way a library function gets reused across a codebase.

Comparison

These three ideas sit at different layers and are complementary, not competing — a common source of confusion worth resolving directly:

ConceptWhat it actually standardizesLayer
Model Context Protocol (MCP)How a model reaches external tools and dataAgent-to-tool connectivity
Agent SDK / frameworkThe loop, state, handoffs, and guardrails around a modelOrchestration and application scaffolding
Multi-Agent System (as a pattern)The architectural design of how multiple agents divide laborConceptual/architectural, implemented using an SDK’s handoff features
Plain API calls, no frameworkNothing beyond what the provider’s raw endpoint doesThe foundation everything else is built on top of

An SDK’s handoff feature is, concretely, one way of implementing the orchestrator-worker or peer-to-peer topologies Multi-Agent System describes abstractly; MCP is what a tool call inside any of those agents typically resolves to when it needs an external system rather than a function defined in the same codebase.

Real-World Use Cases

  • Customer-support triage systems built on an SDK’s handoff primitive, routing between billing, technical, and account-security specialist agents based on the incoming request.
  • Coding agents built on a coding-focused agent SDK, connected to MCP servers for the filesystem, git, and CI systems they need to operate on.
  • Enterprise research assistants combining an SDK’s guardrails with an internal MCP server exposing the company’s own knowledge base.
  • Voice-driven personal assistants using an SDK’s session and state management to hold continuity across a multi-turn spoken conversation, rather than treating each utterance as stateless.
  • Cross-company workflow agents using A2A alongside each participant’s own vendor SDK, so a travel-booking agent from one company can delegate a payment step to a completely separate payments agent built by another.

Common Pitfalls

  • Treating the SDK as a substitute for understanding the underlying mechanics. Debugging a framework’s behavior is much harder without understanding the request/response loop described in Function Calling (Tool Use) that it’s built on top of.
  • Reaching for multi-agent handoffs on a task a single well-scoped agent would handle fine — the same over-engineering pitfall Multi-Agent System warns about, now easier to fall into because the framework makes adding another agent one line of code instead of a real design decision.
  • Shipping without guardrails. Skipping validation in an early prototype and never circling back to add it before real users arrive is a common, costly sequencing mistake.
  • Adopting a vendor-specific SDK without weighing lock-in. The most polished option for a given provider’s models is not automatically the right choice if the team expects to need portability later.
  • Assuming “ADK” always names Google’s specific product. Reading it as the generic term when a specific product was actually meant, or vice versa, causes real confusion in conversations and job postings alike.
  • Delaying tracing until an incident forces it. The same observability pitfall Multi-Agent System calls out for hand-rolled systems applies just as much once a framework is involved — the tooling being available by default doesn’t mean it was actually turned on and used.

Example

A team building a customer-support product picks an agent SDK and defines four agents: a triage agent and three specialists — billing, technical, and account-security. The triage agent’s only job is to read the incoming message and decide which specialist should take over, using the SDK’s handoff primitive rather than a hand-written routing prompt buried in application code.

They connect the whole system to an MCP server already built for their CRM, so all four agents gain read access to customer order history without the team writing a single line of CRM-specific integration code. They add one guardrail: any request to change an account’s password or payment method must pass a secondary-verification check before the responsible agent’s tool call is allowed to execute, regardless of how confident the model is.

A customer messages: “I was charged twice for my last order.” The triage agent hands off to the billing specialist, whose trace — visible afterward through the SDK’s built-in tracing — shows exactly which tool it called against the CRM, what it found, and the exact handoff message that routed the conversation there in the first place. When a bug report comes in later claiming the agent issued a refund it shouldn’t have, that trace is what lets the team reconstruct precisely what happened and why, without piecing it together from scattered application logs never designed to answer that question.

Dig deeper