Function Calling (Tool Use)

Function Calling (Tool Use)

Definition: Function calling (also called tool use) is the mechanism by which a Large Language Model (LLM) invokes external functions, APIs, or systems — a calculator, a database query, a search engine, a code interpreter — as a structured step within generating a response, rather than answering purely from its internal weights. The model doesn’t execute anything itself; it emits a structured, machine-parseable request (typically JSON) naming a tool and its arguments, and the surrounding application executes that request and returns the result. This closes the gap between a model that can only describe an action in prose and a system that can actually perform one, and it is the core primitive underneath every modern AI agent.

How It Works

The Tool Schema

Before a conversation starts, the host application declares a set of available tools to the model. Each tool is described with a name, a natural-language description of what it does and when to use it, and a parameter schema (almost always JSON Schema) specifying argument names, types, constraints, and which arguments are required.

  • The name should be unambiguous and unique — two tools named search and search_v2 invite the model to guess rather than choose deliberately.
  • The tool-level description is the single highest-leverage field: it’s the only place the model learns when to reach for the tool, since it never sees the underlying implementation.
  • Parameter types follow standard JSON Schema primitives (string, number, boolean, array, object), plus constraints like enum for a closed set of valid values and required for arguments that must be present before the call is valid.
  • Nested objects and arrays are fully supported, so a single tool call can carry structured data — a create_event tool might take an attendees array of {name, email} objects rather than a flat list of strings.
  • Many implementations support additionalProperties: false (a “strict mode”) to reject any argument not explicitly declared, closing off a class of malformed or injected calls before they ever reach application code.

The Request/Response Loop

Tool use is fundamentally a multi-turn protocol, even when it feels like one exchange to the end user. The model reads the user’s query plus the tool schemas, decides a tool is needed, and instead of writing a final answer it emits a structured call. The host application intercepts that call, executes the real function, and appends the result back into the model’s context as a new message. The model then continues generating — often producing the final natural-language answer, or deciding another tool call is needed.

Every arrow across the model/tool boundary is a full context round trip — the tool result doesn’t magically appear in the model’s “mind,” it gets serialized to text (or structured content blocks) and re-fed as input, which is why verbose tool results have a direct, measurable cost in tokens and latency.

Tool Choice: Auto, Forced, and None

Most APIs expose a tool_choice setting that controls how much discretion the model has over whether to call a tool at all:

  • auto — the default — lets the model decide per turn whether a tool call is warranted or a direct text answer suffices.
  • any / required — forces the model to make some tool call rather than answering in text, useful when the application’s UI can only render structured output (e.g., a form-filling assistant that must always return a submit_form call).
  • A specific tool name — forces exactly one named tool, useful for testing a single tool in isolation or for multi-step flows where the next action is already known by the application, not the model.
  • none — disables tool use entirely for that turn, even if tools are still declared, which lets an application reuse the same tool-equipped system prompt for a turn where only a plain answer is appropriate.
ModeModel’s discretionTypical use
autoFull — decide per turnGeneral-purpose assistants, chat interfaces
any / requiredMust call something, picks whichUIs that can only render structured output
Named toolNone — that tool, this turnDeterministic multi-step flows, isolated testing
noneZero — text onlyReusing a tool-equipped prompt for a plain-answer turn

Long-Running and Asynchronous Tools

Not every tool finishes fast enough to return its result within the same round trip. A video render, a large data export, or a batch analytics job can take minutes or hours — far longer than a model should sit idle waiting for a response. The common pattern splits such a tool into two: a start_job call that kicks off the work and immediately returns a job_id, and a separate check_job_status call the model can invoke later (either on a timer the application manages, or in response to the user asking “is it done yet?”) to poll for completion. This keeps the request/response loop from Section 2 intact — every individual call is still fast and synchronous — while the underlying task runs asynchronously outside of it. Agent frameworks that support this pattern typically also need a way to resume a conversation after the job completes even if the user isn’t actively chatting, which pushes the design toward webhooks or push notifications rather than the model repeatedly polling on its own.

Parallel, Sequential, and Nested Calls

Early function-calling implementations allowed only one tool call per turn, forcing strictly sequential round trips even when two lookups were independent (e.g., “compare the weather in Jakarta and Manila” required two full round trips). Modern APIs support parallel tool calls, where the model emits multiple structured calls in a single turn, the host executes them concurrently, and all results are returned together before the model continues. Sequential calls remain necessary when a later call depends on an earlier result (e.g., look up a user’s account ID, then use that ID to fetch their orders). Nested or chained tool use — where the output of one tool becomes the input to another, several calls deep — is the basis of agentic workflows like “search for the file, read its contents, then edit it.”

Choosing between the two isn’t left entirely to the model — the schema and prompt design steer it: independent lookups phrased as separate questions (“weather in Jakarta and Manila”) tend to trigger parallel calls naturally, while a tool whose description or required arguments obviously depend on a prior result’s output pushes the model toward sequencing calls one at a time even without an explicit instruction to do so.

Tying Results Back to Calls

When a turn contains more than one tool call, the host application and the model both need a way to match each result to the call that produced it — otherwise a batch of three parallel calls returning three results has no way to know which weather reading belongs to Jakarta and which belongs to Manila. Implementations solve this with a call ID: each emitted tool call carries a unique identifier, and the corresponding result message references that same ID rather than relying on the order results happen to arrive in. This matters in practice because tools frequently complete out of order — a database query might return before a slower web-search call finishes — and the conversation history has to stay internally consistent regardless of completion order. Conversation formats typically also tag these messages with a distinct role (commonly tool or function) separate from user and assistant, so downstream code — and the model itself on the next turn — can tell training-relevant dialogue apart from mechanically-inserted tool output.

Streaming and Partial Arguments

When responses stream token-by-token, tool-call arguments arrive as incremental fragments of JSON rather than a single complete object — a location field might be split across several streamed chunks before it’s parseable. Host applications typically buffer these fragments and only attempt to parse (and execute) the call once the model signals the argument block is complete, since executing against half-formed JSON risks running a call with truncated or default-filled values. This buffering also matters for user experience: a streaming UI that wants to show “calling get_weather…” as soon as the tool name is known, before arguments finish streaming, has to track two different completion points within the same call.

What Happens Under the Hood

The model doesn’t have a separate “function-calling mode” bolted on top of text generation — the same next-token prediction process produces both prose and structured calls. What differs is training: the model is fine-tuned (often via supervised examples plus RLHF (Reinforcement Learning from Human Feedback)) on transcripts that pair user intents with correctly formatted tool calls, teaching it when a call is warranted and how to format one. Many serving stacks additionally use constrained decoding (grammar- or schema-constrained sampling) to guarantee the emitted JSON is syntactically valid against the declared schema, even if the model’s raw token probabilities would otherwise produce malformed output. This is why tool-call arguments almost never fail to parse as JSON, even though argument values can still be wrong — constrained decoding guarantees shape, not correctness.

Tool schemas themselves also consume tokens: every tool definition, description included, is serialized into the model’s context on every single turn of the conversation, whether or not it ends up being called. A system with thirty verbosely-described tools can spend thousands of tokens on schema alone before the user’s first message is even read, which is a direct, recurring cost that scales with tool-set size rather than with actual usage.

Static Tool Lists vs. Dynamic Discovery

The examples above assume tools are hardcoded into the application: a fixed list baked into every request. This works well for a small, stable tool set, but doesn’t scale to an assistant that might need access to hundreds of different internal systems across an organization. The Model Context Protocol (MCP) addresses this by letting a model’s host application discover tools at runtime from one or more external servers — the client asks a connected server “what tools do you expose?” and receives schemas dynamically, rather than a developer hand-writing every schema into the prompt ahead of time. This decouples tool implementation (owned by whoever runs the MCP server) from the application using the model, so the same weather server, database connector, or ticketing-system integration can be reused across many different AI products without each one reimplementing it.

Why It Matters

Tool use isn’t a minor API feature bolted onto chat — it’s the change that made “AI agent” a meaningful product category rather than a research demo:

  • Turns predictors into agents. Tool use is the single mechanism that converts a text-completion engine into something that can search, compute, write files, or move money — the foundation of every product marketed as an “AI agent,” including coding assistants like this one.
  • Grounds answers in live, verifiable data. A model’s parametric knowledge is frozen at training time; a tool call can pull today’s stock price, current weather, or a live database row, sidestepping staleness entirely.
  • Offloads work models are bad at. LLMs are unreliable at exact arithmetic, precise date math, and large-scale lookups. Routing those tasks to a calculator, a SQL engine, or a search index converts a probabilistic guess into a deterministic, checkable result.
  • Standardization is accelerating adoption. Protocols like the Model Context Protocol (MCP) formalize tool discovery and invocation so any compliant client can plug into any compliant tool server, reducing the custom integration work every agent previously required.
  • It’s the substrate of multi-step autonomy. Chained tool calls, where a model plans a sequence of actions and reacts to each result, underlie autonomous coding agents, research agents, and operations agents that complete multi-hour tasks with minimal supervision — see Intelligent Agent and Multi-Agent System.
  • Benchmarks now target it directly. Evaluation suites such as the Berkeley Function-Calling Leaderboard and ToolBench specifically score models on tool-selection accuracy and argument correctness, treating it as a first-class capability alongside reasoning and knowledge.
  • It reshapes product surfaces. Voice assistants, IDE copilots, customer-support bots, and spreadsheet copilots all converged on the same underlying pattern — natural-language front end, function-calling middle layer, deterministic systems in the back — because it’s the most reliable way to let users command software in plain language.
  • It creates a natural permission boundary. Because every action a model can take must first be declared as a tool, the tool list itself becomes an explicit, auditable surface for what an AI system is and isn’t allowed to do — a control point that free-form code execution doesn’t offer.
  • It composes with retrieval. Search and document-lookup tools are themselves just a special case of function calling, which is why Retrieval-Augmented Generation (RAG) pipelines are increasingly implemented as a tool the model chooses to invoke rather than a step that always runs first.
  • It enables graceful escalation. A well-designed tool set lets a model attempt the cheap, fast path first (a cached lookup) and only fall back to an expensive tool (a live web search or a code execution sandbox) when the cheap path is insufficient, rather than hard-coding that branching logic outside the model.

Anatomy of a Tool Definition

A tool definition is just structured metadata — the model never sees the function body, only this description. A typical schema for a weather lookup:

{
  "name": "get_weather",
  "description": "Get current weather conditions for a specific location. Use this whenever the user asks about present-moment temperature, precipitation, or conditions. Do not use for forecasts more than 24 hours out.",
  "parameters": {
    "type": "object",
    "properties": {
      "location": {
        "type": "string",
        "description": "City and country, e.g. 'Jakarta, Indonesia'"
      },
      "unit": {
        "type": "string",
        "enum": ["celsius", "fahrenheit"],
        "description": "Temperature unit for the response"
      }
    },
    "required": ["location"]
  }
}

Given the user query “what’s the weather in Jakarta right now, in Celsius?”, the model emits a call matching that schema:

{
  "tool_call": {
    "name": "get_weather",
    "arguments": {
      "location": "Jakarta, Indonesia",
      "unit": "celsius"
    }
  }
}

Three details in this schema do most of the work. The description field on the tool itself tells the model when to reach for it — omit the “do not use for forecasts” clause and the model may misapply it to forecast questions it can’t actually answer. The enum on unit constrains the value space so the host code never has to handle an unexpected string like "deg C". And required tells the model which arguments must be resolved from the conversation (or asked about) before the call is valid — leave location optional and the model may guess a default city rather than asking the user to clarify.

A second, slightly more complex example shows nested structure — a create_calendar_event tool that must be handed an array of structured attendee objects rather than flat strings:

{
  "name": "create_calendar_event",
  "description": "Create a calendar event with a title, time window, and attendee list. Use only after the user has confirmed the exact time.",
  "parameters": {
    "type": "object",
    "properties": {
      "title": { "type": "string" },
      "start_time": { "type": "string", "description": "ISO 8601 timestamp" },
      "end_time": { "type": "string", "description": "ISO 8601 timestamp" },
      "attendees": {
        "type": "array",
        "items": {
          "type": "object",
          "properties": {
            "name": { "type": "string" },
            "email": { "type": "string" }
          },
          "required": ["email"]
        }
      }
    },
    "required": ["title", "start_time", "end_time"]
  }
}

Nesting required both at the top level (a title and time window must exist) and inside the attendees items (each attendee must at least have an email) lets a single schema enforce two different validity rules at two different levels of the structure.

A few practical rules of thumb consistently improve tool-selection accuracy in real schemas:

  • Write descriptions as usage instructions, not documentation. “Use this to check whether a proposed meeting time conflicts with the attendee’s existing calendar” guides the model far better than a passive restatement like “checks calendar conflicts.”
  • Prefer enum over free-text whenever the value space is closed. A status field with three legal values should never be typed as an unconstrained string — every unconstrained field is an opportunity for a malformed or synonymous-but-wrong value to slip through.
  • Keep required argument counts low. A tool that demands eight required arguments forces the model to either have gathered all eight from the conversation already or start asking clarifying questions before it can act at all; splitting such a tool into a simpler primary call plus optional refinements is usually more reliable.
  • Name tools after the action, not the underlying system. get_customer_orders is more reliable than crm_query_v3, because the model is matching natural-language intent to a name, not looking up an internal API reference.
  • Document units and formats explicitly inside the field description. A budget field described only as "type": "number" invites ambiguity over currency and whether the value is per-night or total; stating “USD, total for the stay” in the description removes the guesswork entirely.
  • Prefer enums over booleans when a third state is plausible. A notify boolean can’t later express “notify, but only for urgent updates” without a breaking change; an enum of ["none", "urgent_only", "all"] leaves room to grow without redefining the field’s type.

Agentic Loops: Chaining Multiple Calls

Single tool calls answer single questions. Real agentic behavior comes from looping the request/response cycle until a task is complete, a pattern often called ReAct (Reason + Act): the model reasons about the current state, takes an action (a tool call), observes the result, and repeats. A file-editing agent asked to “fix the failing test” might chain: read the test file, read the source file it exercises, run the test to see the failure, edit the source, run the test again, and only then report success — five or six tool calls chained on each other’s output, with no human in the loop between steps.

This raises problems single-call tool use doesn’t have:

  • Termination conditions. The loop needs a maximum iteration count or an explicit “done” signal, or a model that keeps deciding “one more lookup would help” will loop indefinitely, burning tokens and latency.
  • Context bloat. Each round trip appends the full tool result to the context, so a chain of ten calls returning verbose JSON can silently consume most of the context window before the model ever produces a final answer, which is why many agent frameworks summarize or truncate older tool results.
  • Compounding errors. Because later calls depend on earlier results, an error early in the chain compounds: if the second call’s arguments were built from a hallucinated value in the first result, every subsequent step inherits that mistake.
  • Retry and error-feedback design. When a tool call fails (a timeout, a 404, an invalid argument rejected by the real API), the failure has to be serialized back into the model’s context just like a success would be — a silently dropped error leaves the model with no signal that anything went wrong, and it may simply fabricate a plausible-sounding answer instead of retrying or surfacing the problem to the user.
  • Idempotency of retries. Retrying a failed call is safe for a read-only tool like a search, but retrying a write tool (charge a card, send an email) without checking whether the first attempt actually succeeded can cause a duplicate side effect — production agent frameworks typically require write tools to accept an idempotency key for exactly this reason.

Total wall-clock latency for a strictly sequential chain of nn calls is roughly

Ttotal=∑i=1n(Tmodel,i+Ttool,i)T_{total} = \sum_{i=1}^{n} \left(T_{model,i} + T_{tool,i}\right)

which is why systems that can parallelize independent branches of a chain — rather than forcing strict sequence — see meaningfully lower latency on complex tasks. A chain that could run four independent lookups in parallel but is implemented sequentially pays for four full round trips instead of one.

Reasoning Interleaved with Tool Calls

Models capable of extended, visible reasoning before answering don’t set that reasoning aside when a tool call is involved — the more capable pattern interleaves the two, reasoning about what’s still unknown, issuing a call to resolve it, reasoning over the result, and deciding whether another call is needed, all within the same overall thinking process rather than as a single reasoning block followed by a single detached call. Preserving that reasoning trace across the tool-call boundary — rather than discarding it once a call is made and starting fresh after the result comes back — helps the model stay consistent about why it made a given call when it later has to interpret that call’s result, which matters most in longer agentic chains where losing track of the original goal a few calls in is a real failure mode.

Evaluating Tool-Use Quality

Tool use introduces failure modes plain text generation doesn’t have, so evaluating it requires metrics beyond standard language-quality scores. Production teams and benchmarks typically track several layers at once:

  • Tool-selection accuracy — given a query, did the model choose to call a tool at all, and if so, the correct one out of the available set? A model that calls search_web when get_account_balance was the obviously correct choice fails here even if the JSON it emits is perfectly formed.
  • Argument correctness — even with the right tool chosen, were the argument values themselves right? This is usually harder to score automatically than tool selection, since “right” can mean exact match (a specific product ID) or semantic equivalence ("NYC" versus "New York City" might both be acceptable).
  • Schema validity — does the emitted call parse as valid JSON and satisfy the declared schema (correct types, required fields present, enum values respected)? Constrained decoding pushes this close to 100% on modern systems, so it’s increasingly treated as a baseline rather than a differentiator.
  • End-to-end task success — for multi-step agentic chains, did the entire sequence of calls accomplish the user’s actual goal? This is the metric that matters most in production but is the hardest to automate, since it usually requires either a held-out set of tasks with known correct outcomes or human judgment.
  • Unnecessary or excessive calls — a model that calls a search tool three times to answer a question a single call could have resolved is technically “successful” but wastes latency and cost; efficient tool use is itself a quality dimension, not just correctness.
  • Recovery behavior — when a call fails or returns an unexpected result, does the model retry sensibly, choose a fallback tool, or ask the user for clarification, rather than either stalling or confidently inventing an answer?

Public benchmarks such as the Berkeley Function-Calling Leaderboard and ToolBench standardize these checks across many models and tool definitions, which is part of why tool-use accuracy is now reported as its own capability score alongside general reasoning and knowledge benchmarks, rather than folded into a single aggregate number.

Security and Permission Design

Because tool use is the mechanism through which a model takes real-world action, the design of the tool layer is itself a security boundary, not just a convenience API. A handful of patterns recur across production systems:

  • Principle of least privilege. Grant a model only the tools a given task genuinely requires — a document-summarization assistant has no reason to be handed a delete_file tool, even if the surrounding application also happens to expose one internally.
  • Separate read tools from write tools. Read-only tools (search, lookup, list) are generally safe to let a model call freely under auto tool choice; write tools (send, delete, charge, publish) benefit from an explicit confirmation step between the model’s call and actual execution.
  • Scope credentials per tool, not per session. A tool that queries a customer’s own order history should authenticate with permissions limited to that customer’s data, so a manipulated or hallucinated argument can’t be used to pivot into another customer’s records.
  • Log every call and result. Because tool calls are the model’s only channel to affect external systems, a complete audit trail of what was called, with what arguments, and what was returned is the primary way to reconstruct what an agent actually did after the fact.
  • Rate-limit and cap iteration counts. Beyond preventing runaway agentic loops (see above), rate limits on expensive or destructive tools bound the damage a single misbehaving chain of calls can do before a human notices.
  • Require human confirmation for irreversible actions specifically, not everything. Gating every single tool call behind a confirmation dialog defeats the purpose of automation; the useful line is drawn at irreversibility — a search can run freely, a deletion should pause for a yes.
  • Sandbox execution tools. Tools that run arbitrary code (a Python or shell interpreter) should execute in an isolated environment with no access to secrets, the wider filesystem, or the network beyond what the task strictly requires — treating model-generated code with the same suspicion as untrusted user input, because functionally that’s what it is.
  • Treat tool output as untrusted input on the way back in. A search result or a scraped page returned from a tool can contain adversarial text aimed at the model itself; robust systems clearly delimit tool output from instructions so the model is less likely to treat embedded text as a new command.
  • Require idempotency keys on write tools. A charge_card or send_message tool that accepts an idempotency key lets the application safely retry a call whose result was lost to a timeout, without risking a duplicate charge or a duplicate message if the original attempt actually succeeded.
  • Default to denying, not allowing, unrecognized tools. If a dynamically-discovered tool set (via Model Context Protocol (MCP) or similar) can change at runtime, the host application should whitelist which discovered tools are actually exposed to the model rather than trusting every tool a connected server happens to advertise.
  • Scope and rotate tool credentials separately from model-provider credentials. The API key a weather tool uses to call a third-party service is a different trust boundary from the key that authenticates calls to the model itself, and should be rotated and revoked independently.

Maintaining Tool Definitions Over Time

Tool schemas aren’t static artifacts written once and forgotten — they’re a live interface between a model and production systems, and they need the same discipline as any other API contract:

  • Renaming a parameter silently breaks calls. If a model was fine-tuned or prompted against a schema using city, and the field is later renamed to location without updating the description or re-testing, tool-selection accuracy can quietly degrade even though nothing “crashed” — the model keeps trying the old name or drops the argument entirely.
  • Version tools explicitly when behavior changes. Introducing get_weather_v2 alongside the original, rather than mutating get_weather in place, avoids breaking whatever agentic flows were already tuned against the old contract, and gives a clean deprecation path.
  • Regression-test schema changes before deploying them. A small held-out set of representative queries, replayed against the new schema, catches accuracy drops from a rewritten description or a changed enum before real users hit them.
  • Track cost alongside correctness. Because every declared tool’s schema is resent on every turn regardless of whether it’s used, teams that add tools freely without pruning unused ones pay a growing, invisible token tax — periodically auditing which tools are actually being called in production is as much a cost exercise as a quality one.
  • Deprecate deliberately. Removing a tool a model has learned to rely on (via fine-tuning or extensive few-shot examples) can cause it to hallucinate a call to a tool that no longer exists; a transition period where the old tool still exists but its description steers the model toward the replacement reduces that risk.
  • Watch for drift between the schema and the real system. If the underlying API a tool wraps changes behavior (a field starts returning a different unit, a new required parameter is added upstream) without the tool schema being updated to match, the model keeps calling it exactly as before while the real result quietly becomes wrong or starts erroring — schema and implementation need to be versioned and reviewed together.

Comparison

Function calling is one of several common strategies for connecting a model to information or capability it doesn’t have natively. They’re often combined rather than chosen exclusively.

DimensionFunction CallingPlain PromptingRetrieval-Augmented Generation (RAG)Fine-Tuning
What it addsLive actions and data via external systemsNothing new — relies entirely on the model’s trained knowledgeRelevant documents retrieved from a knowledge baseNew behavior or knowledge baked into model weights
FreshnessAs current as the tool’s data sourceFrozen at training cutoffAs current as the index is updatedFrozen at training time of the fine-tune
DeterminismHigh — arithmetic, lookups, writes are exactLow — model may approximate or hallucinateMedium — retrieval is deterministic, generation over it isn’tLow — still a generative model, just biased toward the fine-tune data
Can take real-world actionsYes (send email, book meeting, run code)NoNo, unless paired with function callingNo, unless paired with function calling
Setup costModerate — schema design and tool implementationLowest — just a promptModerate to high — indexing pipeline requiredHighest — training data curation plus a training run
Failure modeWrong tool chosen, or malformed/hallucinated argumentsConfident, fluent, factually wrong outputIrrelevant or missing documents retrievedOverfit to training examples, or capability drift on unrelated tasks
Typical useWeather, calculations, bookings, code execution, CRUD operationsCreative writing, summarization of provided text, general reasoningQuestion answering over private or large document setsAdapting tone, format, or domain-specific style at the model level
Added latencyOne or more extra round trips per callNone — single generation passOne retrieval round trip, typicallyNone at inference time — cost is paid upfront during training
Who decides to use itThe model, per tool_choice settingN/A — always activeEither the application (always-on) or the model (as a tool)N/A — baked into every response

RAG and function calling are frequently the same mechanism at different granularity: a “search knowledge base” tool exposed via function calling is a RAG pipeline, just one the model decides to invoke rather than one the application always runs upfront. Fine-tuning, by contrast, changes what the model knows how to do by default — it’s the right tool when the desired behavior should apply on every turn without a live schema to consult, such as always responding in a specific structured format.

In practice, production systems rarely pick just one of these four. A typical enterprise assistant fine-tunes a model on the organization’s preferred tone and output format, gives it RAG access to an internal document index for grounded question-answering, and separately exposes function-calling tools for anything that requires a live system — leaving plain prompting as the fallback path for everything else, like casual conversation or requests the other three approaches don’t apply to. Vendors also differ in schema field names for structurally identical concepts — a top-level functions list versus a tools list containing typed function entries, or parameters versus input_schema — but the underlying contract (name, description, JSON Schema for arguments) is consistent enough that porting a tool definition from one provider’s format to another’s is typically a mechanical relabeling exercise, not a redesign.

Real-World Use Cases

Tool use shows up anywhere a natural-language interface needs to reach a deterministic system underneath it — the pattern is identical whether the underlying tool is a weather API or a multi-million-row production database.

  • Coding assistants and IDE copilots call functions to read files, run tests, execute shell commands, and apply diffs — this document itself was written by a model using exactly this pattern.
  • Customer support bots call CRM and order-management APIs to look up account status, issue refunds, or update a shipping address instead of just describing what the user should do manually.
  • Voice assistants (smart speakers, in-car systems) map spoken requests to function calls that control smart-home devices, set timers, or place calls, translating loosely-phrased speech into precisely-typed arguments.
  • Financial and data-analysis agents call SQL or Python execution tools to pull real numbers and compute exact aggregates rather than approximating them in prose, then chart or summarize the result.
  • Travel and booking assistants chain flight-search, hotel-search, and payment-authorization tools to complete multi-step reservations from a single natural-language request.
  • DevOps and on-call agents query monitoring and logging APIs (metrics dashboards, incident trackers) to diagnose an alert and, in supervised setups, trigger a rollback or restart.
  • Calendar and email assistants call scheduling APIs to check availability and create events, and messaging APIs to draft or send replies, often gated behind a user-confirmation step for anything that sends.
  • Enterprise search and knowledge assistants expose internal document stores, wikis, and databases as tools so employees can query systems in plain language instead of learning each system’s own UI.
  • E-commerce shopping agents call product-search, inventory, and checkout APIs to compare options and complete purchases on a user’s behalf.
  • Multi-agent orchestration systems (Multi-Agent System) use function calling as the inter-agent protocol itself — one agent’s “call another agent” action is implemented as a tool call, with the sub-agent’s final answer returned as the tool result.
  • Legal, medical, and compliance research assistants call citation-lookup and case-database tools to ground every claim in a specific, retrievable source rather than an unverifiable paraphrase from memory.
  • Embedded and IoT assistants running on-device or on a hub call local device-control functions directly — locks, thermostats, industrial sensors — often without a network round trip to a remote server at all.

Common Pitfalls

Most production incidents involving tool use trace back to a handful of recurring design mistakes rather than exotic model failures:

  • Tool sprawl. Registering dozens of overlapping tools (“get_weather,” “get_current_weather,” “fetch_forecast”) confuses tool selection — the model has to guess which near-duplicate is correct, and accuracy drops measurably as the tool count grows.
  • Underspecified descriptions. A tool description that doesn’t say when to use it (versus what it technically does) leads to the model calling it in the wrong situations or ignoring it entirely when it would actually help.
  • Trusting arguments blindly. Model-generated arguments are not validated input — a call to a database or shell tool must still be checked, sanitized, and bounded exactly as if a user had typed it, or a hallucinated argument becomes a real destructive action.
  • No sandboxing on high-stakes tools. Exposing “delete_record” or “send_payment” as a directly callable tool without a confirmation step or scoped permissions turns a model mistake into an irreversible one.
  • Swallowing tool errors. If a tool call fails and the failure isn’t fed back into the model’s context as a result, the model has no way to know it needs to retry, use a different tool, or tell the user something went wrong — it may fabricate a plausible-sounding answer instead.
  • Prompt injection via tool results. Content returned by a tool (a scraped webpage, a document, an email body) can contain text engineered to look like instructions; a model that blindly follows directives found inside tool output rather than treating it as inert data is exploitable.
  • Unbounded agentic loops. Without a hard iteration cap or explicit stop condition, a chained tool-use loop can keep calling “one more” tool indefinitely, burning latency and cost without converging on an answer.
  • Context bloat from verbose results. Returning an entire raw API response (thousands of tokens) instead of a trimmed, relevant subset fills the context window fast, especially across multi-step chains, degrading both cost and downstream reasoning quality.
  • Non-idempotent retries on write tools. Automatically retrying a failed “charge card” or “send email” call without checking whether the original attempt actually went through can silently duplicate a real-world side effect.
  • Confusing schema validity with correctness. Constrained decoding guarantees a call parses as well-formed JSON matching the schema, but teams sometimes treat a successfully parsed call as functionally correct, skipping the argument-level validation that would catch a syntactically valid but semantically wrong value, like a plausible-looking but nonexistent account ID.
  • Ambiguous units and formats. Leaving date formats, currencies, or measurement units unconstrained in a schema (rather than fixed via enum or explicit description) produces calls that parse correctly as JSON but carry the wrong unit — a frequent, hard-to-catch source of silently wrong answers.

Example

A user building a trip-planning assistant gives their model three tools: search_flights(origin, destination, date), get_weather(location, date), and search_hotels(city, checkin, checkout, budget). The user asks: “I want to visit Manila next weekend if the weather’s decent — find me a flight and a cheap hotel.” The model first calls get_weather(location="Manila", date="next weekend"), receives a result showing scattered showers but no storms, and reasons that this counts as “decent.” It then issues two parallel calls — search_flights and search_hotels — since neither depends on the other’s output, and the host application executes both against real booking APIs concurrently rather than making the user wait through two sequential round trips.

The flight search returns four options; the hotel search returns a list filtered by the implied budget signal (“cheap”). The model synthesizes both results into a single natural-language recommendation: a specific flight, a specific hotel under a stated nightly rate, and a one-line note about the weather caveat. Nothing in that final answer was generated from the model’s training data — every concrete fact (flight times, prices, forecast) came from a tool result appended to its context moments earlier.

If the user then replies “actually, book the second flight option,” the model calls a fourth tool, book_flight(flight_id). Because this is a write action with real financial consequences, the application layer intercepts the call before execution and surfaces a confirmation prompt to the user rather than booking immediately — the tool schema declared what could be called, but the host application still decides what’s allowed to execute unsupervised. Once the user confirms, the call executes, a booking confirmation is returned as the tool result, and the model reports the confirmation number back in plain language — closing the loop from natural-language intent to a real-world action, which is the entire behavior chain that “tool use” is built to enable.

Suppose instead the book_flight call times out because the airline’s API is temporarily overloaded. The failure is serialized back into the model’s context as a tool result just like a success would be, carrying an error status rather than a confirmation number. Because that failure is visible to the model rather than silently dropped, it can reason about it correctly — telling the user the booking didn’t go through and offering to retry, rather than fabricating a confirmation number that was never actually issued. This is the practical payoff of treating tool errors as first-class results instead of exceptions the model never sees.

Dig deeper