After reading this post you will be able to wire an LLM API endpoint to a real function in your codebase — a weather lookup, a database query, an internal search route — and understand why the model itself never executes anything. You’ll know the exact shape of the request and response payloads, how to validate the arguments the model produces before you trust them, and the failure modes that will bite you in production if you skip validation.

The tutorial is structured as a beginner path followed by an advanced path. The beginner half covers the minimum viable loop: define a tool, send it to the API, receive a tool call, execute it, send the result back. The advanced half covers parallel calls, strict schema enforcement, error handling, and the specific cases where function calling is the wrong tool for the job.


Beginner Section: The Mental Model

What function calling is not

Function calling is frequently described as “the model using tools.” That phrasing creates a wrong mental model. The model never runs your code. It never opens a socket, never hits your database, never calls an external API. What it does is produce a structured JSON object — the name of a function and a set of arguments — that your application code then decides whether and how to execute.

Think of the model as a translator that converts a natural-language request into a typed function signature. Your application remains the executor. This distinction matters because it determines where your security boundaries live. Every tool call is untrusted input until your code validates it.

The four-message loop

Every function-calling exchange, regardless of provider, follows the same four-step loop:

  1. You send the user’s message plus a list of available tool definitions.
  2. The model responds with either plain text or a structured tool call (function name + JSON arguments).
  3. Your code executes the named function locally with the provided arguments and captures the result.
  4. You send the result back to the model as a new message, and the model produces a final natural-language answer grounded in that result.

Steps 3 and 4 can repeat if the model needs to call multiple tools before it can answer. Some agents loop dozens of times. For a first implementation, keep it to a single round trip.

The tool definition schema

Providers differ slightly in the exact key names, but the shape is consistent: a name, a description, and a JSON Schema describing the parameters. Here is the OpenAI-style format, which you’ll see echoed (with minor key changes) in Anthropic, Gemini, and most OpenAI-compatible endpoints:

{
  "type": "function",
  "function": {
    "name": "get_weather",
    "description": "Return the current weather for a named city. Use only when the user asks about weather in a specific location.",
    "parameters": {
      "type": "object",
      "properties": {
        "city": {
          "type": "string",
          "description": "The city name, e.g. 'Lisbon' or 'New York'"
        },
        "unit": {
          "type": "string",
          "enum": ["celsius", "fahrenheit"],
          "description": "Temperature unit. Default to celsius if unspecified."
        }
      },
      "required": ["city"]
    }
  }
}

Two parts of that definition carry most of the weight and are the parts beginners most often get wrong.

The description field on the function itself is not documentation for humans — it is the prompt the model reads to decide when to call the function. A vague description like “gets weather” causes the model to call the tool when the user asks about climate policy, historical temperatures, or the word “weather” inside an unrelated sentence. A precise description that states when to use and when not to use the function reduces spurious calls dramatically.

The enum constraint on unit is the cheapest quality win available. Without it, the model may emit "C", "metric", "Fahrenheit" with a capital F, or invent a third option. With it, the valid outputs are enumerated and you can validate against a known set rather than string-matching guesses.


The Beginner Path: A Working Weather Tool

Setup

Assume you have Python 3.10+ and an API key for a provider that supports tool calling. Install the official client:

pip install openai
export OPENAI_API_KEY="sk-..."

The same pattern applies to Anthropic (pip install anthropic, client.messages.create with a tools array), so if you’re on Claude, map the key names across — tools instead of tool_choice+functions, and tool_use content blocks instead of tool_calls in the response.

Step 1: Define the tool

Start with the schema above. Rename the function to something your codebase can dispatch on — a prefix like external_ is a useful convention so you can filter tool names at validation time.

Step 2: Send the request

from openai import OpenAI
import json

client = OpenAI()

TOOLS = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Return current weather for a named city.",
        "parameters": {
            "type": "object",
            "properties": {
                "city": {"type": "string"},
                "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}
            },
            "required": ["city"]
        }
    }
}]

response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "What's the weather in Lisbon right now?"}],
    tools=TOOLS,
    tool_choice="auto"
)

message = response.choices[0].message
print(message.tool_calls)

Running that produces a response where message.tool_calls is a list containing something like:

[{
  "id": "call_abc123",
  "type": "function",
  "function": {
    "name": "get_weather",
    "arguments": "{\"city\": \"Lisbon\", \"unit\": \"celsius\"}"
  }
}]

Notice that arguments is a string, not a parsed object. That trips up nearly every first implementation. You must call json.loads() on it yourself, and you must be prepared for that parse to fail — models occasionally emit trailing commas, single quotes, or truncated JSON when the argument is long.

Step 3: Execute locally with validation

This is the step beginners rush and production teams spend the most time on. Do not pass the model’s arguments directly into **kwargs on a function. Validate first.

from pydantic import BaseModel, ValidationError, field_validator

class WeatherArgs(BaseModel):
    city: str
    unit: str = "celsius"

    @field_validator("city")
    @classmethod
    def city_not_empty(cls, v: str) -> str:
        if not v.strip() or len(v) > 80:
            raise ValueError("city must be 1-80 chars")
        return v.strip()

    @field_validator("unit")
    @classmethod
    def unit_in_set(cls, v: str) -> str:
        if v not in {"celsius", "fahrenheit"}:
            raise ValueError("unit must be celsius or fahrenheit")
        return v

def execute_tool_call(tool_call) -> str:
    if tool_call.function.name != "get_weather":
        return json.dumps({"error": "unknown_tool"})
    try:
        parsed = json.loads(tool_call.function.arguments)
        args = WeatherArgs(**parsed)
    except (json.JSONDecodeError, ValidationError) as e:
        return json.dumps({"error": "invalid_arguments", "detail": str(e)})
    result = weather_api.lookup(args.city, args.unit)
    return json.dumps(result)

Three things are happening in that function that are not optional in a real system. First, the tool name is checked against an allow-list before dispatch — this prevents a hallucinated function name from reaching a getattr call on your module. Second, the JSON string is parsed inside a try block. Third, the resulting dict is validated by a schema that mirrors the tool definition. The enum in the tool schema guides the model; the validator enforces it. Guidance and enforcement are separate layers.

Step 4: Return the result and get the final answer

tool_call = message.tool_calls[0]
tool_output = execute_tool_call(tool_call)

followup = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[
        {"role": "user", "content": "What's the weather in Lisbon right now?"},
        message,
        {
            "role": "tool",
            "tool_call_id": tool_call.id,
            "content": tool_output
        }
    ],
    tools=TOOLS
)

print(followup.choices[0].message.content)

The tool_call_id is the link between the model’s request and your response. If you omit it or misalign it, the API returns a 400 error on most providers. This is one of the most common integration bugs — the model produced two tool calls, you only answered one, and now the conversation state is inconsistent.

Result

The model now has the weather data in context and produces a natural-language reply like “It’s currently 22°C in Lisbon with clear skies.” The model never contacted any weather service. Your weather_api.lookup did, and you retained full control over rate limits, credentials, and logging.

What this verifies: a single tool call round trip works end to end. Before moving on, test the failure paths — pass a city name with a SQL-like string in it, force a JSON parse error, and confirm your execute_tool_call returns a structured error rather than throwing.


Advanced Section: Parallel Calls, Strict Mode, and Hard Limits

Parallel tool calls

Modern models can emit multiple tool calls in a single response when a request naturally decomposes: “Compare the weather in Lisbon and Reykjavik.” The model returns two entries in tool_calls, and you must execute both and return two role: "tool" messages, each with the matching tool_call_id. If you only return one, the model will frequently stall or hallucinate the missing value.

Handle this by iterating message.tool_calls rather than indexing [0]. The parallel_tool_calls flag (OpenAI) or the equivalent provider setting lets you disable this behavior if you prefer serial execution — useful when your tools write to shared state and ordering matters.

Strict schema enforcement

Several providers now offer a “strict” or “structured output” mode that constrains decoding so the JSON schema is guaranteed to match. In OpenAI’s API it’s a strict: true key on the function definition:

{
  "type": "function",
  "function": {
    "name": "get_weather",
    "description": "Return current weather for a named city.",
    "strict": true,
    "parameters": {
      "type": "object",
      "properties": {
        "city": {"type": "string"},
        "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}
      },
      "required": ["city", "unit"],
      "additionalProperties": false
    }
  }
}

The trade-off: strict mode requires that every property be listed in required and that additionalProperties be false. You cannot have an optional field. This means you either make the model always supply unit (and give it a default in the description), or you drop the field. Strict mode increases latency slightly and can restrict which models are eligible. For high-stakes tools — anything that writes to a database or triggers a payment — the guaranteed-valid-JSON benefit usually outweighs the cost. For exploratory chat, non-strict is often fine.

The parallelization trap

If you have multiple tools whose names are similar, the model may call the wrong one. A classic case: an search_products tool and a search_orders tool with descriptions that both say “search for matching records.” The model picks one at random because the descriptions don’t contain a discriminating signal. Fix this by naming the kind of entity explicitly in each description (“product catalog,” “customer order history”) and, where possible, trimming the number of tools in a single request to the smallest set that covers the task.

When not to use function calling

Three cases where function calling is the wrong choice:

Pure text transformation. If the user asks for a summary, translation, or formatting change, there is no external data to fetch. Function calling adds a round trip, latency, and a failure mode for zero benefit. Use a plain completion.

Small, deterministic dispatch. If your input space is a fixed set of commands you can match with a regex (“translate X to Y”, “email X about Y”), a deterministic parser is faster, cheaper, and immune to hallucinated tool names. Reserve function calling for cases where the intent is open-ended natural language.

Data you can inject upfront. If the answer fits comfortably in the context window and doesn’t change per-request, inject it into the system prompt. Function calling is for dynamic data — real-time prices, user-specific records, anything that can’t be pasted in ahead of time.

Cost and latency math

Each tool call adds one extra model round trip. Typical small-model round trips run in the hundreds of milliseconds to low seconds depending on load, and every round trip bills both input and output tokens. A naive agent that loops five times on a single user turn pays for six completions. Track tool-call count per turn and set a hard cap — three to five is a reasonable starting ceiling for most chat interfaces. If you exceed it regularly, your tool descriptions are probably not steering the model tightly enough, and you are paying for that ambiguity in tokens and seconds.


A Comparison Table: Beginner vs Advanced Decisions

ConcernBeginner defaultAdvanced choice
Schema keySimple descriptiondescription + strict: true where offered
Tool count per requestAll tools, every requestCurated subset relevant to the task
Argument handlingParse and pass throughPydantic / JSON Schema validation layer
Tool call countSingle call, [0] indexingIterate all calls, cap at N per turn
Error handlingLet exceptions propagateReturn {"error": ...} to the model so it can recover
Name policyFree-formPrefixed namespaces, allow-list check before dispatch
Data injectionAlways call the toolInject static data into system prompt

Common Failure Modes to Test For

Run these against your implementation before you consider the integration finished:

  • Hallucinated tool name. The model invents get_weather_forecast which you never defined. Your dispatcher must reject it and return an error the model can read.
  • Truncated JSON. Long arguments sometimes cut off mid-string. Your json.loads will throw; catch it.
  • Wrong-types-but-valid-JSON. {"city": 42} parses fine but violates the schema. Only a validator layer catches this.
  • Missing tool_call_id in the follow-up. The API rejects the message. Log the ID before you send it.
  • Infinite tool loops. The model keeps calling the same tool with the same arguments. Cap loop depth and detect repeats.
  • Prompt injection through tool output. If your tool returns user-generated text that the model then treats as instructions, you have an injection surface. Sanitize or wrap tool output in clear delimiters.

The most important habit to build: treat every tool call as untrusted input, exactly as you would treat an HTTP request body from an unknown client. The model is a helpful collaborator, not a trusted one.


Wrapping Up: The Decision Rule

If your task requires fresh data, user-specific state, or an action outside the model’s context window, function calling is the standard mechanism and the four-message loop above is the pattern to implement. If the task is text-in, text-out with no external dependency, skip the tool layer entirely — you’ll ship faster and with fewer moving parts.

Start with one tool, validate aggressively, and log every call and its arguments. Once a single round trip is reliable, add parallel calls and strict mode. Do not reach for an agent framework until a hand-rolled loop is visibly straining — the abstraction is convenient, but it hides exactly the validation layer you most need to understand.