After reading this post you will be able to wire an LLM API endpoint to a real function in your codebase — a weather lookup, a database query, an internal search route — and understand why the model itself never executes anything. You’ll know the exact shape of the request and response payloads, how to validate the arguments the model produces before you trust them, and the failure modes that will bite you in production if you skip validation.
The tutorial is structured as a beginner path followed by an advanced path. The beginner half covers the minimum viable loop: define a tool, send it to the API, receive a tool call, execute it, send the result back. The advanced half covers parallel calls, strict schema enforcement, error handling, and the specific cases where function calling is the wrong tool for the job.
Beginner Section: The Mental Model
What function calling is not
Function calling is frequently described as “the model using tools.” That phrasing creates a wrong mental model. The model never runs your code. It never opens a socket, never hits your database, never calls an external API. What it does is produce a structured JSON object — the name of a function and a set of arguments — that your application code then decides whether and how to execute.
Think of the model as a translator that converts a natural-language request into a typed function signature. Your application remains the executor. This distinction matters because it determines where your security boundaries live. Every tool call is untrusted input until your code validates it.
The four-message loop
Every function-calling exchange, regardless of provider, follows the same four-step loop:
- You send the user’s message plus a list of available tool definitions.
- The model responds with either plain text or a structured tool call (function name + JSON arguments).
- Your code executes the named function locally with the provided arguments and captures the result.
- You send the result back to the model as a new message, and the model produces a final natural-language answer grounded in that result.
Steps 3 and 4 can repeat if the model needs to call multiple tools before it can answer. Some agents loop dozens of times. For a first implementation, keep it to a single round trip.
The tool definition schema
Providers differ slightly in the exact key names, but the shape is consistent: a name, a description, and a JSON Schema describing the parameters. Here is the OpenAI-style format, which you’ll see echoed (with minor key changes) in Anthropic, Gemini, and most OpenAI-compatible endpoints:
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Return the current weather for a named city. Use only when the user asks about weather in a specific location.",
"parameters": {
"type": "object",
"properties": {
"city": {
"type": "string",
"description": "The city name, e.g. 'Lisbon' or 'New York'"
},
"unit": {
"type": "string",
"enum": ["celsius", "fahrenheit"],
"description": "Temperature unit. Default to celsius if unspecified."
}
},
"required": ["city"]
}
}
}
Two parts of that definition carry most of the weight and are the parts beginners most often get wrong.
The description field on the function itself is not documentation for humans — it is the prompt the model reads to decide when to call the function. A vague description like “gets weather” causes the model to call the tool when the user asks about climate policy, historical temperatures, or the word “weather” inside an unrelated sentence. A precise description that states when to use and when not to use the function reduces spurious calls dramatically.
The enum constraint on unit is the cheapest quality win available. Without it, the model may emit "C", "metric", "Fahrenheit" with a capital F, or invent a third option. With it, the valid outputs are enumerated and you can validate against a known set rather than string-matching guesses.
The Beginner Path: A Working Weather Tool
Setup
Assume you have Python 3.10+ and an API key for a provider that supports tool calling. Install the official client:
pip install openai
export OPENAI_API_KEY="sk-..."
The same pattern applies to Anthropic (pip install anthropic, client.messages.create with a tools array), so if you’re on Claude, map the key names across — tools instead of tool_choice+functions, and tool_use content blocks instead of tool_calls in the response.
Step 1: Define the tool
Start with the schema above. Rename the function to something your codebase can dispatch on — a prefix like external_ is a useful convention so you can filter tool names at validation time.
Step 2: Send the request
from openai import OpenAI
import json
client = OpenAI()
TOOLS = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Return current weather for a named city.",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string"},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}
},
"required": ["city"]
}
}
}]
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "What's the weather in Lisbon right now?"}],
tools=TOOLS,
tool_choice="auto"
)
message = response.choices[0].message
print(message.tool_calls)
Running that produces a response where message.tool_calls is a list containing something like:
[{
"id": "call_abc123",
"type": "function",
"function": {
"name": "get_weather",
"arguments": "{\"city\": \"Lisbon\", \"unit\": \"celsius\"}"
}
}]
Notice that arguments is a string, not a parsed object. That trips up nearly every first implementation. You must call json.loads() on it yourself, and you must be prepared for that parse to fail — models occasionally emit trailing commas, single quotes, or truncated JSON when the argument is long.
Step 3: Execute locally with validation
This is the step beginners rush and production teams spend the most time on. Do not pass the model’s arguments directly into **kwargs on a function. Validate first.
from pydantic import BaseModel, ValidationError, field_validator
class WeatherArgs(BaseModel):
city: str
unit: str = "celsius"
@field_validator("city")
@classmethod
def city_not_empty(cls, v: str) -> str:
if not v.strip() or len(v) > 80:
raise ValueError("city must be 1-80 chars")
return v.strip()
@field_validator("unit")
@classmethod
def unit_in_set(cls, v: str) -> str:
if v not in {"celsius", "fahrenheit"}:
raise ValueError("unit must be celsius or fahrenheit")
return v
def execute_tool_call(tool_call) -> str:
if tool_call.function.name != "get_weather":
return json.dumps({"error": "unknown_tool"})
try:
parsed = json.loads(tool_call.function.arguments)
args = WeatherArgs(**parsed)
except (json.JSONDecodeError, ValidationError) as e:
return json.dumps({"error": "invalid_arguments", "detail": str(e)})
result = weather_api.lookup(args.city, args.unit)
return json.dumps(result)
Three things are happening in that function that are not optional in a real system. First, the tool name is checked against an allow-list before dispatch — this prevents a hallucinated function name from reaching a getattr call on your module. Second, the JSON string is parsed inside a try block. Third, the resulting dict is validated by a schema that mirrors the tool definition. The enum in the tool schema guides the model; the validator enforces it. Guidance and enforcement are separate layers.
Step 4: Return the result and get the final answer
tool_call = message.tool_calls[0]
tool_output = execute_tool_call(tool_call)
followup = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "user", "content": "What's the weather in Lisbon right now?"},
message,
{
"role": "tool",
"tool_call_id": tool_call.id,
"content": tool_output
}
],
tools=TOOLS
)
print(followup.choices[0].message.content)
The tool_call_id is the link between the model’s request and your response. If you omit it or misalign it, the API returns a 400 error on most providers. This is one of the most common integration bugs — the model produced two tool calls, you only answered one, and now the conversation state is inconsistent.
Result
The model now has the weather data in context and produces a natural-language reply like “It’s currently 22°C in Lisbon with clear skies.” The model never contacted any weather service. Your weather_api.lookup did, and you retained full control over rate limits, credentials, and logging.
What this verifies: a single tool call round trip works end to end. Before moving on, test the failure paths — pass a city name with a SQL-like string in it, force a JSON parse error, and confirm your execute_tool_call returns a structured error rather than throwing.
Advanced Section: Parallel Calls, Strict Mode, and Hard Limits
Parallel tool calls
Modern models can emit multiple tool calls in a single response when a request naturally decomposes: “Compare the weather in Lisbon and Reykjavik.” The model returns two entries in tool_calls, and you must execute both and return two role: "tool" messages, each with the matching tool_call_id. If you only return one, the model will frequently stall or hallucinate the missing value.
Handle this by iterating message.tool_calls rather than indexing [0]. The parallel_tool_calls flag (OpenAI) or the equivalent provider setting lets you disable this behavior if you prefer serial execution — useful when your tools write to shared state and ordering matters.
Strict schema enforcement
Several providers now offer a “strict” or “structured output” mode that constrains decoding so the JSON schema is guaranteed to match. In OpenAI’s API it’s a strict: true key on the function definition:
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Return current weather for a named city.",
"strict": true,
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string"},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}
},
"required": ["city", "unit"],
"additionalProperties": false
}
}
}
The trade-off: strict mode requires that every property be listed in required and that additionalProperties be false. You cannot have an optional field. This means you either make the model always supply unit (and give it a default in the description), or you drop the field. Strict mode increases latency slightly and can restrict which models are eligible. For high-stakes tools — anything that writes to a database or triggers a payment — the guaranteed-valid-JSON benefit usually outweighs the cost. For exploratory chat, non-strict is often fine.
The parallelization trap
If you have multiple tools whose names are similar, the model may call the wrong one. A classic case: an search_products tool and a search_orders tool with descriptions that both say “search for matching records.” The model picks one at random because the descriptions don’t contain a discriminating signal. Fix this by naming the kind of entity explicitly in each description (“product catalog,” “customer order history”) and, where possible, trimming the number of tools in a single request to the smallest set that covers the task.
When not to use function calling
Three cases where function calling is the wrong choice:
Pure text transformation. If the user asks for a summary, translation, or formatting change, there is no external data to fetch. Function calling adds a round trip, latency, and a failure mode for zero benefit. Use a plain completion.
Small, deterministic dispatch. If your input space is a fixed set of commands you can match with a regex (“translate X to Y”, “email X about Y”), a deterministic parser is faster, cheaper, and immune to hallucinated tool names. Reserve function calling for cases where the intent is open-ended natural language.
Data you can inject upfront. If the answer fits comfortably in the context window and doesn’t change per-request, inject it into the system prompt. Function calling is for dynamic data — real-time prices, user-specific records, anything that can’t be pasted in ahead of time.
Cost and latency math
Each tool call adds one extra model round trip. Typical small-model round trips run in the hundreds of milliseconds to low seconds depending on load, and every round trip bills both input and output tokens. A naive agent that loops five times on a single user turn pays for six completions. Track tool-call count per turn and set a hard cap — three to five is a reasonable starting ceiling for most chat interfaces. If you exceed it regularly, your tool descriptions are probably not steering the model tightly enough, and you are paying for that ambiguity in tokens and seconds.
A Comparison Table: Beginner vs Advanced Decisions
| Concern | Beginner default | Advanced choice |
|---|---|---|
| Schema key | Simple description | description + strict: true where offered |
| Tool count per request | All tools, every request | Curated subset relevant to the task |
| Argument handling | Parse and pass through | Pydantic / JSON Schema validation layer |
| Tool call count | Single call, [0] indexing | Iterate all calls, cap at N per turn |
| Error handling | Let exceptions propagate | Return {"error": ...} to the model so it can recover |
| Name policy | Free-form | Prefixed namespaces, allow-list check before dispatch |
| Data injection | Always call the tool | Inject static data into system prompt |
Common Failure Modes to Test For
Run these against your implementation before you consider the integration finished:
- Hallucinated tool name. The model invents
get_weather_forecastwhich you never defined. Your dispatcher must reject it and return an error the model can read. - Truncated JSON. Long arguments sometimes cut off mid-string. Your
json.loadswill throw; catch it. - Wrong-types-but-valid-JSON.
{"city": 42}parses fine but violates the schema. Only a validator layer catches this. - Missing
tool_call_idin the follow-up. The API rejects the message. Log the ID before you send it. - Infinite tool loops. The model keeps calling the same tool with the same arguments. Cap loop depth and detect repeats.
- Prompt injection through tool output. If your tool returns user-generated text that the model then treats as instructions, you have an injection surface. Sanitize or wrap tool output in clear delimiters.
The most important habit to build: treat every tool call as untrusted input, exactly as you would treat an HTTP request body from an unknown client. The model is a helpful collaborator, not a trusted one.
Wrapping Up: The Decision Rule
If your task requires fresh data, user-specific state, or an action outside the model’s context window, function calling is the standard mechanism and the four-message loop above is the pattern to implement. If the task is text-in, text-out with no external dependency, skip the tool layer entirely — you’ll ship faster and with fewer moving parts.
Start with one tool, validate aggressively, and log every call and its arguments. Once a single round trip is reliable, add parallel calls and strict mode. Do not reach for an agent framework until a hand-rolled loop is visibly straining — the abstraction is convenient, but it hides exactly the validation layer you most need to understand.
🔗 Recommended Reading
- How to Evaluate LLM Outputs: A Beginner's Step-by-Step Tutorial to Testing and Scoring AI Responses
- LLM Guardrails for Beginners: A Step-by-Step Tutorial to Filtering Unsafe Outputs
- Building Your First AI Agent: A Beginner Tutorial with Python and the OpenAI API
- RAG Chunking Strategies for Beginners: A Step-by-Step Tutorial
- Prompt Engineering Mistakes That Break JSON Output and How to Fix Them