Two things developers new to LLM integration routinely conflate are “the model calls my API” and “the model tells me it wants me to call my API.” These are not the same operation, and mistaking one for the other is the source of most confusing behavior in early function-calling projects. In practice, the model never touches your network stack, your credentials, or your database. It emits a structured request — a JSON object naming a function and supplying arguments — and your application decides whether and how to execute that request. Everything else in this tutorial follows from that single architectural fact.
This post uses a myth-vs-reality structure, because the gap between what function calling sounds like and what it does is where beginners lose the most time. Each section pairs a common misconception with the way the mechanism works, and then shows the concrete code or behavior that reconciles them.
Myth 1: “The model executes functions.”
Reality: The model produces a tool-call payload, and your code executes it.
This distinction matters more than any other in the topic. A language model is a text generator. It has no socket to your weather API, no ability to run curl, and no awareness that a network exists. When you hand it a tool definition, you are telling it: “If it helps you answer, emit a structured object in this shape and I will handle the rest.”
The API response typically contains a tool_calls array on the assistant message. Each entry has an id, a type (usually "function"), and a function object with a name and a arguments string containing JSON. Your application parses arguments, invokes the matching local handler, and then sends the result back as a new message with role tool (OpenAI-style) or as a tool_result content block (Anthropic-style). The model sees that result and continues.
Here is what a tool call looks like in a raw OpenAI chat completion response:
{
"id": "chatcmpl-9aX2f...",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": null,
"tool_calls": [
{
"id": "call_abc123",
"type": "function",
"function": {
"name": "get_current_weather",
"arguments": "{\"location\":\"Berlin, DE\",\"unit\":\"celsius\"}"
}
}
]
},
"finish_reason": "tool_calls"
}
]
}
Notice that arguments is a JSON string, not a nested object. A very common beginner bug is treating it as an object and reading call.function.arguments.location directly. That returns undefined. You have to JSON.parse it first, and you have to wrap that parse in a try/catch, because models occasionally emit malformed JSON when the schema is loose or the parameter descriptions are vague.
Myth 2: “I just describe the function in the system prompt.”
Reality: Tools are declared through a dedicated schema, separate from the prompt.
You can describe your API in prose in a system prompt and hope the model responds with valid JSON, but that path is fragile. The supported mechanism is a tools array where each entry has a type of "function" and a function object containing name, description, and parameters (a JSON Schema object). The model is trained to look at this array and produce matching output. You get validation on the provider side, and you get consistent tool_calls structure back.
A complete, working tool definition for a weather lookup looks like this:
{
"type": "function",
"function": {
"name": "get_current_weather",
"description": "Get the current weather for a given city. Use this whenever the user asks about current conditions, temperature, or precipitation in a specific location.",
"parameters": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "City and country, e.g. 'Berlin, DE' or 'Austin, TX'."
},
"unit": {
"type": "string",
"enum": ["celsius", "fahrenheit"],
"description": "Temperature unit. Default to celsius unless the user specifies otherwise."
}
},
"required": ["location"]
}
}
}
Three details in that schema do most of the reliability work. First, the description on the function itself tells the model when to use it, not just what it does — that “Use this whenever…” clause reduces spurious calls. Second, the enum on unit constrains the model to two valid values, which eliminates a whole class of garbage arguments. Third, required is set explicitly; without it, the model may omit location and force your handler to guess.
If your schema omits a description on a parameter, expect the model to infer reasonable-but-wrong values. Parameter descriptions are not documentation for humans — they are the primary signal the model uses to fill the field.
Myth 3: “Function calling means the model picks the tool and that’s the end.”
Reality: It is a loop. The model may call multiple tools, chain results, or request nothing at all.
Modern providers support parallel tool calls: in one response, the model can emit three tool_calls entries because it decided it needs weather in three cities. Your code must iterate over the array, execute each one, and return all results in a single follow-up message. If you only handle the first entry, the model will see an incomplete picture and may call again, wasting a round trip.
The loop terminates when finish_reason is "stop" (or "tool_calls" no longer appears) and the assistant message contains prose content instead of a tool call. Your code should also cap iterations — commonly 5 to 10 — to prevent runaway loops where the model keeps calling a tool that returns an error it cannot resolve.
Here’s a concrete implementation path that wires this together with the OpenAI Python SDK. This assumes OPENAI_API_KEY is set in your environment.
import json
import os
from openai import OpenAI
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
def get_current_weather(location: str, unit: str = "celsius") -> dict:
# Replace with a real call to your weather provider.
# This stub keeps the tutorial runnable without external credentials.
return {"location": location, "unit": unit, "temp": 18, "condition": "clear"}
TOOLS = [{
"type": "function",
"function": {
"name": "get_current_weather",
"description": "Get current weather for a city. Use when the user asks about weather or temperature.",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string", "description": "City and country, e.g. 'Berlin, DE'."},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}
},
"required": ["location"]
}
}
}]
def run_conversation(user_message: str, max_turns: int = 6):
messages = [{"role": "user", "content": user_message}]
for _ in range(max_turns):
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=messages,
tools=TOOLS,
tool_choice="auto",
)
msg = response.choices[0].message
messages.append(msg)
if not msg.tool_calls:
return msg.content
for call in msg.tool_calls:
try:
args = json.loads(call.function.arguments)
except json.JSONDecodeError:
args = {}
if call.function.name == "get_current_weather":
result = get_current_weather(**args)
else:
result = {"error": f"unknown tool: {call.function.name}"}
messages.append({
"role": "tool",
"tool_call_id": call.id,
"content": json.dumps(result),
})
return "Max tool iterations reached without a final answer."
print(run_conversation("What's the weather in Berlin and Austin right now?"))
Two things in that code are easy to get wrong. First, messages.append(msg) appends the entire assistant message object, including the tool_calls array — that is required so the follow-up tool messages have something to attach to via tool_call_id. If you construct a plain {"role": "assistant", "content": msg.content} instead, the API rejects the next request. Second, the tool result must be a string; passing a dict will raise a validation error. json.dumps(result) is the standard fix.
Run that script and you should see something like: “Right now it’s 18°C and clear in Berlin, and 18°C and clear in Austin.” — assuming the stub returns the same values. Swap the stub for a real HTTP call and the loop is production-shaped.
Myth 4: “Once it works, I don’t need to think about failure modes.”
Reality: The interesting bugs all live at the boundaries — parse errors, schema drift, tool-choice ambiguity, and infinite loops.
Four failure modes dominate beginner and intermediate integrations:
Malformed arguments. Models occasionally emit {"location": "Berlin" with a missing brace when the schema is complex. Always wrap JSON.parse / json.loads in a try/except and decide on a fallback: either skip the call and report an error to the model (which lets it retry), or abort. Providers are trending toward “strict mode” tool schemas that constrain generation and make this rarer, but it is not eliminated.
Unknown tool names. If you list a tool in tools but forget to handle it in the dispatch switch, the model can call it and your code returns an error. Return the error as a tool message rather than throwing — the model can often course-correct when it sees the failure text.
Tool-choice ambiguity. With tool_choice="auto" (the default), the model decides whether to call a tool or answer directly. If your tool’s description overlaps conceptually with another tool, the model may pick the wrong one. Tighter descriptions and non-overlapping names fix most cases. If you need to force a specific call on the first turn — common in structured-extraction pipelines — set tool_choice={"type": "function", "function": {"name": "your_tool"}}.
Runaway loops. If a tool keeps returning an error the model cannot reason past, it may retry the same call repeatedly. The max_turns counter in the example above is the standard mitigation. Without it, a single bad request can consume your rate limit and your budget.
Myth 5: “More tools = more capable.”
Reality: Every tool you add competes for the model’s attention and increases mis-selection risk.
There is a widely observed pattern where tool selection accuracy degrades as the declared tool count grows. The mechanism is the same as context dilution elsewhere in prompting: more options mean more plausible-but-wrong matches. Systems that need dozens of tools typically split them into named sets and route per-request, presenting the model with the 3–8 tools relevant to the current task rather than the full catalog. This is sometimes called “tool retrieval” and it mirrors how RAG narrows context before generation.
Concretely: if you build an agent with 40 tools, expect to invest in routing, not just schemas. A simple first step is to organize tools into groups (e.g., billing_*, calendar_*, support_*) and select the group based on the user’s first message or a lightweight classification step.
Myth 6: “Function calling is always the right answer.”
Reality: For many tasks, a plain structured-output request is simpler, cheaper, and less error-prone.
If your goal is to extract fields from a text — pull a date, an amount, and a vendor name out of an invoice — you do not need a tool. You need a prompt that asks for a JSON object matching a schema, ideally with the provider’s structured-output feature (OpenAI’s response_format={"type": "json_schema", ...}, Anthropic’s tool-use with forced tool choice). Function calling exists to let the model request actions it cannot perform itself: querying live data, writing to a database, hitting a third-party API, or running a calculation it would otherwise guess at.
The rule of thumb: use function calling when the model needs a side effect or fresh data. Use structured output when you just want the shape of the response constrained.
A second trade-off is latency and cost. Each tool round trip is a full model inference. A request that chains three tools is roughly three times the token cost of a single-turn answer, and users feel the delay. If a task can be done with one API call and no model loop, do that instead. Reserve the loop for cases where the decision of which API to call depends on natural-language reasoning.
A minimal debugging checklist
When a function-calling integration misbehaves, work through this order:
- Log the raw
tool_callspayload. Ifargumentsis not valid JSON, the schema is too loose — addenum,required, and tighter parameter descriptions. - Confirm you are appending the full assistant message to
messages, not a stripped-down object.tool_call_idmismatches are the most common validation failure. - Verify tool result content is a string, not a dict.
- Check
finish_reason. If it is"tool_calls"and you returned a final answer anyway, you cut the loop short. - Add a turn counter and cap it. Unbounded loops are the most expensive beginner bug.
- If the model keeps picking the wrong tool, shorten the list before you rewrite descriptions.
None of these require a different model or a heavier framework. They are the same debugging instincts you would apply to any RPC boundary: validate the payload, respect the protocol, and bound the retries.
The next time a request to an LLM-tool pipeline fails, resist the urge to rewrite the whole prompt. Inspect the tool-call payload first — more often than not, the fix is one tightened schema field or one missing tool_call_id, not a redesign.
🔗 Recommended Reading
- Vector Databases Explained: A Beginner's Guide to Storing and Querying Embeddings
- How to Evaluate LLM Outputs: A Beginner's Step-by-Step Tutorial to Testing and Scoring AI Responses
- LLM Guardrails for Beginners: A Step-by-Step Tutorial to Filtering Unsafe Outputs
- Building Your First AI Agent: A Beginner Tutorial with Python and the OpenAI API
- Function Calling in LLM APIs: A Beginner Tutorial for Connecting AI to Real Tools