An LLM call and an AI agent are not the same thing, and the gap between them is where most beginner tutorials lose people. A single API call takes a prompt, returns text, and stops. An agent takes a goal, decides which actions to take, executes them, observes the results, and keeps looping until the goal is met or it gives up. The model is the same in both cases. What changes is the control flow wrapped around it.
That difference matters because it determines what you have to build. If you want a chatbot, you write a loop that appends messages to a list and re-sends them. If you want an agent, you write a loop that appends messages, inspects the response for tool calls, executes those tools, feeds the results back, and repeats — with a stopping condition so it doesn’t run forever. The tutorial below walks through the second case using Python and the OpenAI API, from a bare tool-calling call up to a running agent with two tools and a hard iteration cap.
Myth: An agent is a smarter model
Reality: An agent is a control loop around an ordinary model, plus a tool schema, plus a stopping rule.
The model itself has no memory between calls and no ability to take actions on its own. Every capability you associate with “agentic” behavior — browsing, running code, querying a database, editing a file — is a function you wrote, registered with the model as a callable, and executed by your code when the model asks for it. The model’s only job during a tool call is to emit a structured JSON object saying “call this function with these arguments.” Your program decides whether to honor that request, runs the function, and hands the result back.
This is why the same model can behave like a passive chatbot in one application and an autonomous problem-solver in another. The intelligence is comparable; the scaffolding is not.
Myth: You need a framework like LangChain or CrewAI to build an agent
Reality: Frameworks help with multi-agent orchestration and long-running state, but a functional single-agent system fits in under 100 lines of plain Python.
The OpenAI Python SDK, or any equivalent client for Anthropic, Google, or a local model, exposes everything you need: chat completions with tools, structured tool-call objects in the response, and the ability to append tool results as new messages. Frameworks layer abstractions on top of that, and abstractions are worth adopting when they solve a problem you have. For a first agent, they mostly obscure the mechanism you’re trying to learn.
Build the loop by hand once. You’ll understand every framework you pick up afterward, and you’ll know exactly which part of the abstraction is doing the work.
Myth: The prompt is the hard part
Reality: The tool schema and the stopping condition cause more agent failures than the system prompt.
A vague system prompt produces vague behavior, yes. But a tool whose description is ambiguous, whose parameters are untyped, or whose failure modes are silent will break an agent that otherwise has a perfect prompt. The same is true of the loop: without a maximum iteration count and a clear terminal state, an agent can spin on the same failed tool call indefinitely, burning tokens and time.
Tool design and loop control are engineering problems. Prompt engineering is a smaller piece of the puzzle than tutorials suggest.
Myth: Tool calling means the model executes code
Reality: The model only requests a call. Execution happens in your process, under your control.
When the API returns a tool_calls array, nothing has run yet. The payload is a suggestion: a function name and a JSON string of arguments. Your code parses that string, looks up the function in a registry you defined, runs it, and appends the result as a message with role="tool". If you don’t want to run it — because the arguments look wrong, or because the tool is dangerous — you simply don’t, and you return an error message instead.
That gate is the entire security model of a tool-calling agent. It’s also the reason agents are not inherently dangerous: the model can request anything, but it can only affect what your code is willing to execute.
Setting up the project
Before writing any agent logic, get a working environment. The steps are short.
# Create an isolated environment
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
# Install the SDK
pip install openai
# Store the key as an environment variable
export OPENAI_API_KEY="sk-..."
On Windows, set the key with set OPENAI_API_KEY=sk-... in cmd or $env:OPENAI_API_KEY="sk-..." in PowerShell. Never hardcode the key in the source file. If the file is committed, the key is leaked.
The SDK reads OPENAI_API_KEY from the environment automatically when you instantiate the client with no arguments. That’s the configuration path to prefer.
A single tool call, stripped to its essentials
Start smaller than an agent. Write one function, register it with the model, and handle the tool call manually. This is the primitive that the agent loop repeats.
import json
import os
from openai import OpenAI
client = OpenAI()
# 1. Define a plain Python function. This is what runs on your machine.
def get_weather(city: str) -> str:
"""Return a mock weather report for a given city."""
fake_db = {
"tokyo": "18C, light rain",
"berlin": "11C, overcast",
"austin": "29C, clear",
}
return fake_db.get(city.lower(), f"No data for {city}")
# 2. Describe the tool to the model using JSON Schema.
tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city. Use this whenever the user asks about weather conditions.",
"parameters": {
"type": "object",
"properties": {
"city": {
"type": "string",
"description": "The city name, in lowercase, e.g. 'tokyo'.",
}
},
"required": ["city"],
},
},
}
]
# 3. Make the call. Note that we do NOT get a normal text answer back.
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "What's the weather in Berlin?"}],
tools=tools,
)
message = response.choices[0].message
# 4. Inspect the tool call. The model has not run anything yet.
if message.tool_calls:
call = message.tool_calls[0]
args = json.loads(call.function.arguments) # arguments arrive as a JSON string
result = get_weather(**args)
print(f"Model wants: {call.function.name}({args})")
print(f"Your code returned: {result}")
else:
print("Model answered directly:", message.content)
Run that and you’ll see output like Model wants: get_weather({'city': 'berlin'}) followed by the weather string. The model requested a call; your code executed it. Nothing in the API executed anything.
Two details tend to trip beginners here. First, call.function.arguments is a string, not a dict. It needs json.loads. Second, the function name in the schema and the Python function name must match if you’re doing a registry lookup by name — this is where typos produce silent no-ops.
Closing the loop: from tool call to agent
A single call plus a tool execution is a tool-calling system, not an agent. The jump to an agent is small: append the model’s tool-call message to the conversation, append the tool’s result, call the model again, and repeat until the model stops requesting tools. That repeated cycle is the agent loop.
import json
from openai import OpenAI
client = OpenAI()
def get_weather(city: str) -> str:
fake_db = {"tokyo": "18C, light rain", "berlin": "11C, overcast", "austin": "29C, clear"}
return fake_db.get(city.lower(), f"No data for {city}")
def convert_temp(celsius: float, to: str = "fahrenheit") -> str:
"""Convert a Celsius temperature to Fahrenheit or Kelvin."""
if to == "fahrenheit":
return f"{(celsius * 9 / 5) + 32:.1f}F"
if to == "kelvin":
return f"{celsius + 273.15:.1f}K"
return "Unsupported unit"
# Tool registry: name -> callable. This is how the model's request maps to code.
TOOL_REGISTRY = {
"get_weather": get_weather,
"convert_temp": convert_temp,
}
TOOLS = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather for a city. Call this when the user asks about weather.",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "City name, lowercase."}
},
"required": ["city"],
},
},
},
{
"type": "function",
"function": {
"name": "convert_temp",
"description": "Convert a temperature from Celsius to Fahrenheit or Kelvin.",
"parameters": {
"type": "object",
"properties": {
"celsius": {"type": "number"},
"to": {"type": "string", "enum": ["fahrenheit", "kelvin"]},
},
"required": ["celsius", "to"],
},
},
},
]
SYSTEM_PROMPT = (
"You are a helpful weather assistant. When the user asks a question, "
"call the appropriate tool. If a tool call fails, explain the error and stop."
)
def run_agent(user_input: str, max_iterations: int = 6) -> str:
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": user_input},
]
for step in range(max_iterations):
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=messages,
tools=TOOLS,
)
message = response.choices[0].message
# No tool call -> the model is done. Return the final text.
if not message.tool_calls:
return message.content or "(empty response)"
# Append the assistant's tool-call message so the model remembers what it asked for.
messages.append(message)
# Execute every requested tool and append each result.
for call in message.tool_calls:
name = call.function.name
try:
args = json.loads(call.function.arguments)
fn = TOOL_REGISTRY[name]
result = fn(**args)
except KeyError:
result = f"ERROR: unknown tool '{name}'"
except json.JSONDecodeError:
result = "ERROR: could not parse tool arguments as JSON"
except TypeError as e:
result = f"ERROR: bad arguments for {name}: {e}"
messages.append({
"role": "tool",
"tool_call_id": call.id,
"content": str(result),
})
return "ERROR: agent exceeded max iterations without producing a final answer"
if __name__ == "__main__":
print(run_agent("What's the weather in Austin, and what is that in Fahrenheit?"))
That script is a complete agent. It has a goal, tools, a loop, a stopping condition, and error handling on each tool call. Run it, and the exchange looks roughly like this:
- User asks about Austin’s weather and its Fahrenheit equivalent.
- Model requests
get_weather(city="austin"). - Your code runs it, appends
"29C, clear"as a tool message. - Model now requests
convert_temp(celsius=29, to="fahrenheit"). - Your code runs it, appends
"84.2F". - Model returns a plain text answer and the loop exits.
That’s the pattern. Everything an agent framework does for you is a variation on it.
The failure modes that will bite you first
Infinite tool loops. If a tool always returns the same unhelpful result, the model will often retry with a slightly different argument forever. The max_iterations cap prevents this. Six is a reasonable default for tutorials; production agents sometimes need more, but they should also have a smarter exit condition than “counter reached N.”
Unparseable arguments. The arguments arrive as a string. Sometimes it’s valid JSON, sometimes it isn’t, and sometimes the model invents a parameter your function doesn’t accept. Wrap json.loads and the function call in try/except and return a descriptive error string. The model reads that string on the next iteration and usually corrects course. A tool that throws an uncaught exception crashes your whole agent.
Schema drift between description and function. If your tool description says “city name” but your function expects “location”, the model will pass the wrong key and hit a TypeError. Keep the JSON Schema and the Python signature aligned.
Tool result bloat. A tool that returns 50KB of raw JSON will fill the context window in a few iterations and start costing real money. Truncate, summarize, or return only the fields the model needs.
Silent tool failures. A tool that returns an empty string on failure, with no explanation, gives the model nothing to work with. Prefer "ERROR: <reason>" over "" or None.
Hallucinated tool names. If the model has been trained on common tool names like search or browse and you haven’t registered one, it may still request it. The registry lookup fails with KeyError, which you should catch and return as a tool error so the model can pick a registered tool on the next pass.
When not to build an agent
Not every task needs a loop. Skip the agent architecture when:
- One API call solves the problem. If a well-formed prompt reliably produces the right answer, wrapping it in a tool-calling loop adds latency, cost, and failure surface for no benefit.
- The tool is a single deterministic function. If you know the exact function to call from the user’s input, just call it. The model doesn’t need to decide.
- You need predictable latency. Agent loops run the model multiple times per user request. Latency for a two-step agent is often two to four times a single call. If your application has a strict response-time budget, an agent is the wrong shape.
- You can’t afford nondeterminism. The same input can produce different tool sequences across runs. If you need a deterministic pipeline, write one.
- The cost per user request is sensitive. Every iteration is a full model call. Token spend scales with iteration count, and iteration count is not fully under your control.
Agents are the right tool when the task requires multiple steps whose order depends on intermediate results, and when a small amount of nondeterminism is acceptable.
Extending the agent from here
Once the loop runs, the improvements that matter most are usually in the tools and the stopping logic, not the prompt.
- Add a
list_toolsor capability description tool so the model can ask what it can do when it’s unsure. Useful when the tool count grows past a handful. - Log every message and tool call to a file. Debugging an agent without a transcript is painful. A JSON-lines log per session costs nothing and pays back constantly.
- Add a
final_answertool. Instead of parsing the model’s last text message, give it a dedicated tool that explicitly terminates the loop with a structured answer. This makes the stopping condition unambiguous. - Constrain argument types with enums and patterns. The more the schema constrains valid inputs, the fewer parse errors you see.
- Consider parallelism. If a step requires two independent tool calls, you can execute them concurrently before appending results. This reduces latency meaningfully on multi-tool turns.
None of these require a framework. They’re small edits to the same loop you just wrote.
A quick comparison of design choices
| Decision | Simple choice | When to escalate |
|---|---|---|
| Model | Small, cheap model | Switch when tool selection accuracy drops |
| Tools | 1 to 3 well-described functions | Add a capability discovery tool past ~8 tools |
| Loop stop | max_iterations counter | Add explicit final_answer tool for structure |
| Errors | Return “ERROR: reason” string to model | Log to file and alert on repeated failure |
| Memory | In-memory message list | Persist to disk or DB for multi-session agents |
| Parsing args | json.loads with try/except | Consider Pydantic models for validation |
The table is a rough ordering of effort versus return, not a rule. Most first agents work fine with the left column.
What to verify before calling the agent done
Three checks catch the majority of beginner bugs:
- Force a tool call. Ask a question the tools must handle and confirm the loop runs at least two iterations.
- Force an error. Temporarily rename a tool in the registry and confirm the agent returns a tool error and either recovers or terminates gracefully. If it crashes, your exception handling is incomplete.
- Force a runaway. Ask a question with no possible tool answer and confirm the agent hits
max_iterationsrather than looping forever.
If all three pass, the agent is structurally sound. Prompt tuning and tool refinement come after that, and they’re incremental rather than foundational.
The clearest way to think about the whole exercise: you are not building an intelligent system, you are building a reliable one that happens to be driven by an intelligent component. The reliability comes from your tools, your loop, and your error handling. The intelligence comes from a model that is, in the end, just another function call.
🔗 Recommended Reading
- How to Evaluate LLM Outputs: A Beginner's Step-by-Step Tutorial to Testing and Scoring AI Responses
- LLM Guardrails for Beginners: A Step-by-Step Tutorial to Filtering Unsafe Outputs
- Function Calling in LLM APIs: A Beginner Tutorial for Connecting AI to Real Tools
- RAG Chunking Strategies for Beginners: A Step-by-Step Tutorial
- Prompt Engineering Mistakes That Break JSON Output and How to Fix Them