Model Context Protocol (MCP)

Model Context Protocol (MCP)

Definition: The Model Context Protocol (MCP) is an open protocol, originally released by Anthropic, that standardizes how AI applications connect to external tools, data, and prompt templates. Rather than every application hand-writing bespoke integration code for every system it wants a model to reach — one company’s Slack integration, another’s database connector — MCP defines a common client-server contract: an MCP server exposes a system’s capabilities in a standard shape, and any MCP client embedded in a host application (a chat app, an IDE, a coding agent) can discover and use them the same way, regardless of which model or vendor built the host. It’s often described as “a USB-C port for AI applications” — one connector shape that any compliant device on either side can plug into, replacing a tangle of one-off cables. MCP doesn’t replace Function Calling (Tool Use) — a tool call is still how a model actually invokes something — it standardizes how that tool got discovered and reached in the first place.

How It Works

Host, Client, and Server Roles

MCP defines three distinct roles, and keeping them straight is the key to understanding everything else about the protocol:

  • The host is the user-facing AI application — a desktop assistant, an IDE, a custom agent — that owns the connection to the language model and decides what the model is allowed to see and do.
  • The client lives inside the host and maintains a single, stateful, one-to-one connection to exactly one server. A host that wants to reach three different systems runs three clients internally, one per server.
  • The server is a separate, typically small program that exposes one system’s capabilities — a GitHub connector, a filesystem, a company’s internal ticketing tool — over the protocol. A server has no idea which host or model is calling it; it just answers protocol requests.

A single host commonly runs many clients at once, one per connected server, and aggregates everything they expose into one combined capability set the model can draw on for that conversation.

The Three Core Primitives

MCP servers expose capability through three distinct primitives, each controlled by a different actor:

  • Tools are callable functions — the model decides, at inference time, whether and when to invoke one, exactly like a function call in the underlying Function Calling (Tool Use) mechanism. Tools are model-controlled.
  • Resources are read-only data — a file, a database row, a document — that the host application can attach to a conversation’s context. The model doesn’t request a resource itself the way it calls a tool; the host decides what to attach and when. Resources are application-controlled.
  • Prompts are reusable, parameterized prompt templates a server defines and a host can surface directly to a human user, often as a slash command or menu item — “summarize this PR,” “draft a release note.” Prompts are user-controlled.
PrimitiveWho decides to use itAnalogous to
ToolThe model, per-turnA function call
ResourceThe host applicationAn attached file or context document
PromptThe human userA saved template or slash command

This three-way split matters when designing a server: something the model should reach for autonomously belongs in tools; something that should always be available as background context belongs in resources; a canned workflow a person explicitly triggers belongs in prompts.

Transport Layer

Every MCP message is JSON-RPC 2.0, regardless of how it physically travels. Two transports cover most real deployments:

TransportHow it worksBest for
stdioThe host launches the server as a local subprocess and exchanges messages over its stdin/stdoutLocal dev tools — a filesystem server, a local git server — with zero network exposure
Streamable HTTPThe server runs as an independent, possibly remote, process; clients connect over HTTP and can receive streamed responsesRemote or shared servers, multiple simultaneous clients, anything needing real authentication

stdio is the simpler and more private option — nothing the server touches ever leaves the machine unless the tool itself makes an outbound call — which is why most reference implementations for sensitive local resources (a filesystem, a local database) default to it. HTTP-based transport trades that isolation for reach: a company can run one MCP server centrally and let every employee’s AI tool connect to it over the network.

Capability Negotiation

Before any real exchange happens, the client and server perform an initialize handshake: each side states its protocol version and which capabilities it supports, so a client never assumes a server offers something it doesn’t. Once initialized, the client can enumerate what’s on offer — tools/list, resources/list, prompts/list — and later invoke them with tools/call, resources/read, or prompts/get.

That final call-and-result exchange is exactly the request/response loop described in Function Calling (Tool Use) — MCP hasn’t changed what happens at the model boundary, only how the tool’s existence and schema were discovered in the first place, and over what channel the call actually travels.

MCP vs. Hardcoded Tool Lists

Without MCP, every application wanting a “search GitHub issues” capability has to write and maintain that integration itself — the schema, the API calls, the auth, the error handling — independently, once per application. With MCP, one GitHub server implements that capability a single time, and any MCP-compliant host can connect to it and gain the exact same capability with zero reimplementation. This reframes integration from an M × N problem — M applications, each needing custom glue for N systems — into an M + N problem: each system builds one server, each application builds one client, and the two sides connect freely without either needing to know about the other in advance.

Why It Matters

  • Decouples tool implementation from tool consumption. A GitHub server, a database connector, or a ticketing-system integration gets built once and reused across every compliant AI product, instead of being rewritten inside each one.
  • Turns an M×N integration problem into M+N. This is the single biggest structural reason MCP spread as fast as it did — it changes the shape of the integration cost curve, not just who pays it.
  • Enables genuine runtime tool discovery. A host doesn’t need to know in advance what a server offers; it asks, and the answer can change over time as the server’s own capabilities evolve.
  • Keeps sensitive data local when it needs to. The stdio transport means a server touching private files or an internal database never has to expose anything over a network at all.
  • Standardization is accelerating adoption across vendors. Multiple major AI providers and countless third-party tools now speak MCP, which is precisely the point of an open protocol — no single vendor’s AI product is the only thing that can use a given server.
  • It composes with, rather than replaces, the underlying mechanism. MCP tools still resolve to ordinary model tool calls at the point of execution — see Function Calling (Tool Use) — so everything already understood about tool-call reliability, argument correctness, and context cost still applies.
  • It complements agent-to-agent protocols rather than covering the same ground. MCP connects a model to tools and data; a separate class of protocol (such as Agent2Agent, or A2A — see Agent SDKs and Frameworks) connects one autonomous agent to another. The two problems look similar but aren’t the same one.
  • It creates an explicit, auditable trust boundary. Because every server a host connects to is a distinct, nameable thing, deciding what to trust becomes a concrete question — “should this host connect to this server” — rather than an implicit property buried inside application code.

Security Considerations

Because an MCP server is, functionally, a piece of third-party code the host is choosing to trust, connecting to one carries real risk that’s specific to this protocol, not just tool-use risk in general:

  • Tool poisoning. A malicious or compromised server can describe a tool dishonestly — a description that claims to do something safe while the underlying implementation does something else — manipulating the model into calling it under false pretenses.
  • The confused-deputy problem. A host with legitimate, high-privilege access to one sensitive server can be tricked — often via adversarial content returned from a second, less-trusted server — into misusing that first access on the attacker’s behalf.
  • Rug-pull risk. A server’s declared tools can change between the moment a host lists them and the moment it calls one, silently altering behavior the host developer never explicitly approved.
  • Prompt injection via resources. Content pulled in through a resource — a file, a scraped page, a database row — can carry adversarial instructions the model may treat as commands unless the host clearly delimits data from instructions, the same general risk Function Calling (Tool Use) describes for tool results, now arriving through a second primitive as well.
  • Supply-chain trust. Connecting to a third-party MCP server is functionally similar to adding a new dependency to a codebase — it deserves the same scrutiny (who publishes it, what does it actually do, what does it have access to) before being granted a host’s tools or data.

The practical mitigations mirror general tool-use security: prefer local stdio servers for anything sensitive, review a server’s declared tools before connecting, scope what each server is allowed to touch, and treat everything a server returns as untrusted input rather than a trusted instruction.

Comparison

ApproachDiscoveryCross-vendor reuseSetup costRuns where
MCPDynamic, at runtime, via the protocolHigh — any compliant host can use any compliant serverModerate — one server implementation, reusable everywhereLocal (stdio) or remote (HTTP)
Hardcoded Function Calling (Tool Use)Static — a fixed list baked into the applicationNone — the integration lives inside one application’s codeLow per integration, but repeats for every applicationWherever the host application runs
A REST API integrated directlyNone — the application calls it explicitly, no model-facing discoveryLow — each integration is bespokeModerate, and repeats per applicationWherever the API is hosted
Browser/chat-app plugin systems (an earlier analogous idea)Semi-dynamic, but scoped to one vendor’s platformLow — plugins built for one platform don’t transfer to anotherModerateVendor-hosted

MCP’s real distinction from hardcoded function calling isn’t the model-facing mechanics — a called MCP tool and a hardcoded tool look identical to the model — it’s that the schema and implementation live outside the application entirely, discoverable and reusable by anything else that speaks the same protocol.

Real-World Use Cases

  • Coding agents and IDEs connecting to a filesystem server, a git server, and a database server to read, edit, and query a project without any of that integration code living inside the editor itself.
  • Enterprise knowledge assistants connecting to a company’s internal ticketing system, wiki, or CRM, each exposed as one server that every AI tool across the organization can share.
  • Desktop AI applications connecting to local tools over stdio — a local file server, a local git server — keeping sensitive data entirely on-device.
  • Personal productivity agents connecting a calendar, an email client, and a note-taking app, each as its own server, composed together by whichever host the user happens to prefer.
  • Data analysis agents connecting to a database server that exposes safe, pre-scoped query tools instead of raw, unrestricted SQL access.
  • Design-to-code workflows where a coding agent pulls layout and asset context directly from a design tool’s server mid-session, instead of a developer manually re-describing it.

Common Pitfalls

  • Trusting every discovered server by default. A server appearing in a tools/list response is not the same as a server being safe to grant real access to — trust has to be a deliberate decision, not an implicit one.
  • Overexposing write permissions. A server that only needs to read a calendar shouldn’t also be able to delete events, for the same least-privilege reasoning that applies to any tool — see Function Calling (Tool Use).
  • Connecting to too many servers at once. Every connected server’s tools are serialized into the model’s context, so a host aggregating a dozen servers can pay a large, recurring token cost before the conversation even starts.
  • Confusing MCP with agent-to-agent protocols. MCP connects a model to tools and data; it says nothing about how two independent agents should talk to each other — that’s a separate problem, covered under Agent SDKs and Frameworks.
  • Assuming a server’s capabilities are static. A server can legitimately change what it offers between sessions; code that hardcodes assumptions about a specific server’s tool list rather than re-checking it can silently break.
  • Skipping the initialize handshake’s capability check. Assuming a server supports a feature without confirming it during negotiation risks a request the server has no way to answer correctly.
  • Running a sensitive local server over a network transport out of convenience. Anything that doesn’t need to leave the machine shouldn’t — stdio exists specifically for this case.

Example

A team is building an internal coding agent and wants it to search their company’s ticket tracker and read from their private GitHub organization. Before MCP, this meant writing and maintaining two bespoke integrations inside the agent’s own codebase — one per system, each with its own auth handling and schema, neither reusable if the team later built a second AI tool that needed the same access.

Instead, they connect their agent host to two existing MCP servers: one already built for their ticket tracker, one already built for GitHub. On startup, the host launches an MCP client for each, and both perform the initialize handshake before anything else happens. The host then calls tools/list on each server and learns that the ticketing server exposes search_issues and create_issue, while the GitHub server exposes search_code and get_file — schemas the team never had to write by hand.

A developer asks the agent, “is there an open bug about the login timeout, and if so, what does the relevant code look like?” The model decides a search is needed, and the host’s client sends tools/call search_issues(query="login timeout") to the ticketing server, which returns a matching ticket. The model then calls search_code on the GitHub server with a term drawn from that ticket’s description, gets back a matching file, and synthesizes both results into one answer — a bug summary plus the relevant code — without the coding agent’s own codebase containing a single line of ticket-tracker or GitHub-specific integration logic. If the team ships a second AI tool next quarter, both servers are already built and ready to connect to, at zero additional integration cost.

Dig deeper