Tool Use and Function Calling
How an LLM asks your code to run functions: tool definitions, JSON Schema inputs, tool_use and tool_result blocks.
What is it?
Tool use (also called function calling) lets a model ask your application to run a function and give it the result. The model cannot browse, read your database or do reliable arithmetic on its own, but it can say: 'please call get_weather with {"city": "Paris"}'. Your code runs it and replies with the answer. This single mechanism is what turns a text generator into something that can look things up and take actions.
A tool definition has three parts:
- name: a short identifier like
get_weather(letters, digits, underscores, hyphens). - description: plain-language text telling the model what the tool does, when to use it, and what it returns. This is the most important part - the model chooses tools mostly by reading descriptions.
- input_schema: a JSON Schema - a standard JSON format for describing the shape of data - that lists the arguments, their types, which are required, and allowed values (
enum).
The conversation then contains two new kinds of content blocks (a message's content can be a list of typed blocks instead of a plain string):
- tool_use (in the assistant's message):
{type: "tool_use", id: "toolu_...", name: "get_weather", input: {city: "Paris"}}. Theidis unique per call. - tool_result (in the next user message):
{type: "tool_result", tool_use_id: "toolu_...", content: "18°C and cloudy"}. Thetool_use_idmust match the call it answers. Addis_error: truewhen the tool failed, with an error message as the content, so the model knows to try something else.
When the model wants a tool, the response's stop_reason is tool_use. You must append the assistant message exactly as received (including the tool_use blocks) to your history, run the tools, then send one user message containing the tool_result blocks. Then call the model again.
Parallel tool calls: a single assistant turn can contain several tool_use blocks (for example, weather in Paris and Oslo). Run them all (in parallel if they are independent) and return all results in one user message, one tool_result per tool_use id. Splitting them across several messages is a common bug.
tool_choice controls whether the model may use tools. The default, auto, lets the model decide whether to call a tool or answer directly. On current Claude models, rely on auto plus clear instructions in the system prompt and tool descriptions ('Always use the calculator for arithmetic') rather than trying to force a specific tool.
Writing good tools is a design skill, sometimes called the agent-computer interface:
- Write descriptions like documentation for a new colleague: what it does, when to use it, when not to, what the output looks like, limits (max rows, units).
- Keep the set small and non-overlapping. Two tools that both 'search' confuse the model.
- Make arguments hard to misuse: enums instead of free text, clear units ('amount_cents'), absolute or workspace-relative paths stated explicitly.
- Return concise, meaningful output. A 2 MB JSON dump wastes context; return the fields that matter, and say when results were truncated.
- Return helpful error messages ('No file named todo.txt; available: notes.txt, plan.md') so the model can recover.
Tool use is also a convenient way to get structured data: the model's input object always follows the shape you described. For pure data extraction, dedicated structured-output features (see Structured Output) are usually simpler.
Explain like I'm 10
Think of a manager on the phone with an assistant at the office. The manager cannot reach the filing cabinet, so they say: 'Please pull the file for order 42.' The assistant fetches it and reads it back. The manager decides what to ask next. The manager is the model, the assistant is your code, and the list of things the assistant is allowed to do - with instructions for each - is your tool list.
Examples
Validate and dispatch tool calls like an API would send them (runnable)
// One tool definition, in the same shape you send to the API.
const tools = [{
name: "get_weather",
description: "Get the current temperature for a city. Use when the user asks about weather.",
input_schema: {
type: "object",
properties: {
city: { type: "string", description: "City name, e.g. Paris" },
unit: { type: "string", enum: ["celsius", "fahrenheit"], description: "Defaults to celsius" },
},
required: ["city"],
},
}];
// A tiny JSON Schema check (real projects use a library such as ajv or pydantic).
function validate(schema, input) {
const errors = [];
for (const key of schema.required || []) if (!(key in input)) errors.push("missing required field '" + key + "'");
for (const [key, value] of Object.entries(input)) {
const prop = schema.properties[key];
if (!prop) { errors.push("unknown field '" + key + "'"); continue; }
if (prop.type === "string" && typeof value !== "string") errors.push(key + " must be a string");
if (prop.enum && !prop.enum.includes(value)) errors.push(key + " must be one of: " + prop.enum.join(", "));
}
return errors;
}
const impl = {
get_weather: ({ city, unit = "celsius" }) => {
const c = { paris: 18, oslo: 4 }[city.toLowerCase()];
if (c === undefined) throw new Error("No weather data for " + city + ". Known cities: Paris, Oslo");
return unit === "celsius" ? c + " C" : Math.round(c * 9 / 5 + 32) + " F";
},
};
// Turn ONE tool_use block into ONE tool_result block - never throw.
function handleToolUse(block) {
const def = tools.find((t) => t.name === block.name);
if (!def) return { type: "tool_result", tool_use_id: block.id, content: "Unknown tool: " + block.name, is_error: true };
const errors = validate(def.input_schema, block.input);
if (errors.length) return { type: "tool_result", tool_use_id: block.id, content: "Invalid input: " + errors.join("; "), is_error: true };
try {
return { type: "tool_result", tool_use_id: block.id, content: impl[block.name](block.input) };
} catch (err) {
return { type: "tool_result", tool_use_id: block.id, content: err.message, is_error: true };
}
}
// Pretend the model asked for four tools in ONE turn (parallel tool calls).
const assistantContent = [
{ type: "text", text: "Let me check both cities." },
{ type: "tool_use", id: "toolu_1", name: "get_weather", input: { city: "Paris" } },
{ type: "tool_use", id: "toolu_2", name: "get_weather", input: { city: "Oslo", unit: "fahrenheit" } },
{ type: "tool_use", id: "toolu_3", name: "get_weather", input: { town: "Rome" } },
{ type: "tool_use", id: "toolu_4", name: "get_stock_price", input: {} },
];
const results = assistantContent.filter((b) => b.type === "tool_use").map(handleToolUse);
// ALL results go back in ONE user message:
const nextUserMessage = { role: "user", content: results };
for (const r of nextUserMessage.content) console.log(JSON.stringify(r));Every tool_use gets exactly one tool_result with the matching id, errors are reported (not thrown) with is_error so the model can recover, and all results travel together in one user message.
One full tool round trip with the Anthropic SDK
import anthropic
client = anthropic.Anthropic()
tools = [{
"name": "get_weather",
"description": "Get the current weather for a city. Use when the user asks about weather. "
"Returns temperature in Celsius and a short condition, e.g. '18 C, cloudy'.",
"input_schema": {
"type": "object",
"properties": {"city": {"type": "string", "description": "City name, e.g. 'Paris'"}},
"required": ["city"],
},
}]
def get_weather(city: str) -> str:
fake = {"paris": "18 C, cloudy", "oslo": "4 C, snow"}
if city.lower() not in fake:
raise ValueError(f"No data for {city}. Try a major city.")
return fake[city.lower()]
messages = [{"role": "user", "content": "Should I pack a coat for Paris and Oslo?"}]
# Turn 1: the model asks for tools
resp = client.messages.create(model="claude-opus-5-5", max_tokens=16000,
tools=tools, messages=messages)
print(resp.stop_reason) # "tool_use"
messages.append({"role": "assistant", "content": resp.content}) # keep tool_use blocks!
results = []
for block in resp.content:
if block.type == "tool_use":
try:
out = get_weather(**block.input)
results.append({"type": "tool_result", "tool_use_id": block.id, "content": out})
except Exception as e:
results.append({"type": "tool_result", "tool_use_id": block.id,
"content": f"Error: {e}", "is_error": True})
messages.append({"role": "user", "content": results}) # ALL results, ONE message
# Turn 2: the model uses the results to answer
resp = client.messages.create(model="claude-opus-5-5", max_tokens=16000,
tools=tools, messages=messages)
print(next(b.text for b in resp.content if b.type == "text"))Two calls: the first returns tool_use blocks, the second receives the results. Turning these two calls into a loop gives you an agent (see The Agent Loop).
What the message history looks like after one tool round
[
{"role": "user", "content": "Should I pack a coat for Paris and Oslo?"},
{"role": "assistant", "content": [
{"type": "text", "text": "I'll check the weather in both cities."},
{"type": "tool_use", "id": "toolu_01A", "name": "get_weather", "input": {"city": "Paris"}},
{"type": "tool_use", "id": "toolu_01B", "name": "get_weather", "input": {"city": "Oslo"}}
]},
{"role": "user", "content": [
{"type": "tool_result", "tool_use_id": "toolu_01A", "content": "18 C, cloudy"},
{"type": "tool_result", "tool_use_id": "toolu_01B", "content": "4 C, snow"}
]}
]Roles still alternate user/assistant. Tool results are sent with role 'user' because they come from your side of the conversation. The ids shown are illustrative.
How it works
When you pass tools, the API turns your definitions into special instructions that the model was trained to follow. The model reads the conversation and the tool descriptions, and if it decides a tool would help, it generates a tool_use block instead of (or after) plain text and stops with stop_reason: "tool_use".
The model generates the input object to match your JSON Schema, but treat it as untrusted input: validate types and ranges, check permissions, and never pass it straight to a shell or SQL query. A model can be tricked (by text in a web page or document it read) into requesting harmful actions - this is prompt injection, covered in AI Security.
Tool definitions are sent with every request and count as input tokens, so a list of 50 verbose tools costs on every call. Keep the list focused on the task.
your app model
-------- -----
messages + tool defs ------> reads, decides
<------ tool_use {id, name, input}
validate input (stop_reason = tool_use)
run real function
tool_result {tool_use_id,
content, is_error?} ------> reads result
<------ final text
(stop_reason = end_turn)Why does it exist?
Models are frozen at training time and live inside a text box. They do not know today's weather, your customer's order status, or the contents of your repository, and they make arithmetic slips. Tool use gives them a disciplined, structured way to ask for fresh data and to trigger actions, while your code keeps control over what actually runs.
When to use it
Use tools whenever the model needs information it cannot know (live data, private data, search), needs exact computation (math, date arithmetic, code execution), or needs to take an action (create a ticket, write a file). Also use a tool definition when you want guaranteed-shape arguments from the model.
When not to use it
Do not wrap things in tools that you could just put in the prompt: if the model always needs the same small document, include it directly. Do not expose dangerous capabilities (arbitrary shell, raw SQL, payments) as tools without strict validation and approval. Do not create dozens of overlapping tools; merge or remove them.
Common mistakes
Writing one-line tool descriptions; the model picks tools by their descriptions, so vague ones cause wrong or missing calls.
Dropping the assistant's tool_use blocks from the history and sending only the text - the next request then has tool_results with no matching tool_use and fails.
Sending parallel tool results as separate user messages instead of one message with several tool_result blocks.
Raising an exception and crashing the program when a tool fails, instead of returning a tool_result with is_error: true.
Trusting tool inputs blindly - passing a model-supplied path or query straight to the file system, shell or database.
Returning huge raw outputs (whole web pages, full tables) that flood the context window.
Defining overlapping tools (search_docs and find_documents) so the model has to guess which one to use.
Practice exercises
- Easy:
Write a tool definition (name, description, input_schema) for
convert_currencywith amount, from and to fields, using enums for currency codes. - Easy:
In the runnable demo, add a
daysinteger field to get_weather and extend the validator to check numbers. Send an invalid value and confirm the error result. - Medium:
Rewrite three tool descriptions from a real API you know so that a new colleague could use them correctly without other docs. Include when NOT to use each.
- Medium:
Using the Python round-trip example, ask a question that needs no tool ('What is the capital of France?') and print the stop_reason. Explain the result.
- Hard:
Build a dispatcher in Python that runs independent tool calls from one turn concurrently with a thread pool, preserves the matching ids, applies a 10-second timeout per tool, and reports timeouts as is_error results.
Interview questions
Walk through one tool-use round trip.
You send messages plus tool definitions. The model returns an assistant message with one or more tool_use blocks and stop_reason tool_use. You append that assistant message unchanged, run each tool, then append one user message with a tool_result per tool_use id. You call the model again; it reads the results and either calls more tools or answers.
Why does the tool description matter so much?
The model decides whether and which tool to call almost entirely from names and descriptions. A good description states what the tool does, when to use it and not use it, argument meanings and units, and what the output looks like. Most tool-selection bugs are description bugs.
How should a tool failure be reported to the model?
Return a tool_result for that tool_use id with is_error set to true and a clear, actionable message (what failed and how to fix it). The model can then retry with different arguments or choose another approach instead of the program crashing.
How do you handle parallel tool calls?
Run every tool_use in the assistant turn (concurrently if independent) and send all tool_result blocks together in a single user message, each referencing its tool_use_id.
Is a tool's input safe to use directly?
No. It is model-generated and can be influenced by prompt injection from content the model read. Validate against the schema, enforce permissions and allow-lists, sanitize paths and queries, and require human approval for risky actions.
What does tool_choice auto mean?
The model decides on its own whether to call a tool or answer directly. On current Claude models you guide this with clear system-prompt instructions and descriptions rather than forcing a specific tool.