Stage 3 — Tool Use & Your First Agent Loop ⭐¶
This stage does one thing: let the model fill out a “tool work order,” then have your program validate it, execute it, and send the result back. This round trip is your first Agent Loop.
📌 Learning Objectives¶
By the end, you can:
- Name the five steps:
schema → call → execute → result → answer. - Define a tool, validate its arguments, and safely run the corresponding function.
- Write an Agent Loop with a step limit and a stopping condition, without a framework.
- Tell Function Calling and Structured Output apart instead of treating them as the same thing.
- Compare schemas or models with fixed prompts rather than drawing a conclusion from one result.
🚪 Entry Conditions¶
If you can run a Python file, understand functions and dicts, and have completed Stage 02, you are ready. If your environment is not ready, go back to Stage 00 first.
🧩 Eight Core Terms First¶
Tool Use¶
When a model needs external data or an action, it first makes a tool request. It is like a child asking an adult to open a box on a high shelf: the model says what it wants done, and the program actually acts. This chapter uses it to check weather and do calculations. The model itself does not execute your client tool.
Function Calling¶
The model returns a function name and arguments in an agreed format. It is like filling out a work order with fixed fields. This chapter uses it to turn a natural-language question into a request that a program can read. Message formats are not identical across providers.
Tool Schema¶
JSON (JavaScript Object Notation) is a text format for sharing data.
A schema is a tool’s information card: its name, purpose, fields, and data types. It is like a menu telling a customer what can be ordered. This chapter describes tools with JSON Schema. A schema constrains the shape, but the program must still validate values, permissions, and business rules.
Tool Call¶
A Tool Call is the work order filled out by the model. It contains the tool name, call ID, and arguments. For example: “Check Taipei, using Celsius.” This chapter’s program reads it first, then finds an allowed function in the allowlist. It is a request, not an execution result.
Tool Result¶
A Tool Result is the data returned after the program finishes the work, matched back to the original request by call ID. It is like a kitchen putting a completed dish on the correct table. This chapter sends successful or error results back to the model. External results may be untrusted and must not be treated as highest-priority instructions.
Agent Loop¶
The program repeats “ask the model → execute a tool → return the result” until it gets an answer or reaches a limit. It is like following a recipe one step at a time and stopping when it is done. The full round trip is model → tool call → execute → tool result → model. This chapter’s working definition is model + tools + bounded loop; it is a learning definition, not the only academic definition of an Agent.
ReAct¶
ReAct alternates between deciding the next step, taking an action, observing what happened, and continuing. It is like looking on the table for your keys first, then checking a drawer if they are not there. The loop written here is ReAct-inspired and observable; it does not require the model to reveal private Chain-of-Thought.
Structured Output¶
The model returns data in a fixed shape, such as JSON that conforms to a schema. It is like filling an answer into a form. This chapter contrasts it with Function Calling: the former asks for data, while the latter asks a program to take an action. Even a valid shape can contain wrong content, a refusal, or truncated output.
Choose the Right Method First¶
| What you need | Start with | Example |
|---|---|---|
| Only a text answer | A normal model response | Rewrite an email |
| Data in a fixed shape | Structured Output | Extract a name and date |
| Live data or an action | Function Calling / Tool Use | Check weather, create a ticket |
⚠️ Five Guardrails Before Writing Your First Agent¶
- Execute only tools in the allowlist; never use a model-generated name for arbitrary function calls.
- Treat tool arguments as untrusted input; validate types, ranges, and permissions first.
- Give a tool only the minimum permissions needed to complete the task.
- Require human confirmation before high-risk actions such as deleting, paying, or sending email.
- Set a maximum number of turns, a timeout, and a cost limit; do not let the Agent loop forever.
📚 Required Reading¶
Read in this order:
- Ollama Tool Calling ⭐⭐⭐⭐⭐ — Start with the single-tool and multi-turn loop.
- Anthropic — How Tool Use Works ⭐⭐⭐⭐⭐ — See what the model, application, and tool result each do.
- ReAct paper ⭐⭐⭐⭐ — Read the abstract first; learn where Reasoning + Acting comes from without trying to finish every equation at once.
Expand prerequisites, setup, time, and budget
Prerequisites: You can run Python, understand lists/dicts/functions, and have completed Stage 02.
Primary local path: Ollama + qwen2.5:3b. This is the beginner model retained after verifying the user’s installation; it is not claimed to be best for every schema.
ollama pull qwen2.5:3b
ollama serve
python -m pip install "openai>=3.3,<4"
Cloud comparison path: Anthropic + a pinned Haiku model ID.
$env:ANTHROPIC_API_KEY="paste-your-key-here"
python -m pip install "anthropic>=1.0,<2"
On macOS/Linux, set it with export ANTHROPIC_API_KEY="paste-your-key-here". Do not put the key in code or commit it.
Time: Plan about 2–3 hours for Exercises 1–3, about 3–5 hours for Exercises 4–6, and 5–8 hours for the full active path.
Cost calculation:
cost = input tokens ÷ 1,000,000 × input price
+ output tokens ÷ 1,000,000 × output price
At the 2026-08-27 check, Claude Haiku 4.5 was $1 / $5 (input / output, per million tokens). If one request uses 2,000 input + 1,000 output tokens, the example cost is about $0.007. A tool loop sends multiple requests; reserve $0.05 per exercise first, and set a $1 provider spend limit for five full-chapter experiments. These are conservative caps, not billing guarantees.
Path A has $0 in API cost; it still uses your hardware, memory, and electricity.
Classic Agent Paradigms (thinking patterns)¶
Expand for the differences between CoT, ReAct, Reflection, and Planning
| Term | Plain-language use | Where to learn it |
|---|---|---|
| Chain-of-Thought (CoT) | Early prompt techniques often asked for intermediate reasoning. Do not treat a full private chain of thought as a general output requirement; when checking, look at the final answer and a short, verifiable reason | Stage 02 |
| ReAct | Alternate actions and observations in a loop, then decide the next step | Exercise 3 of this chapter |
| Reflection | A broad practice of using one round of feedback to improve the next attempt | The routing section below |
| Reflexion / Self-Refine | Research patterns with an explicit Actor/Critic or self-feedback process | This chapter’s concepts; persistent-memory version in Stage 06 |
| Planning | Break the task into steps, then adjust the plan from results | Stage 07.5 |
These terms describe different ways to solve problems, not the only test for whether something is an Agent. Computer-use, CodeAct, and workflow agents may use different loops too.
🛠 Hands-on Exercises¶
Complete Exercises 1–3 first. Exercises 4–6 make the loop more robust; you do not need to finish them all in one day.
Exercise 1: Function Calling (One Tool, One Call)¶
After finishing, you will see the model first produce a get_weather Tool Call, the program execute it, and the model answer using the Tool Result.
If you prefer working from files, open the complete Exercise 1 folder.
First action: Copy and run ollama pull qwen2.5:3b. Then expand Path A and copy the complete program into hello_tool.py.
Path A: Complete copyable Ollama example (API cost $0)
import json
from openai import OpenAI
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
TOOLS = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get demonstration weather data for a specified city",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "City name, for example Taipei"},
"unit": {"type": "string", "enum": ["celsius"]},
},
"required": ["city", "unit"],
"additionalProperties": False,
},
},
}]
def get_weather(city: str, unit: str) -> dict:
if unit != "celsius":
raise ValueError("Only celsius is accepted")
return {"city": city, "temperature": 26, "unit": unit}
messages = [{"role": "user", "content": "What is the temperature in Taipei now?"}]
first = client.chat.completions.create(
model="qwen2.5:3b", messages=messages, tools=TOOLS
)
assistant = first.choices[0].message
messages.append(assistant.model_dump(exclude_none=True))
for call in assistant.tool_calls or []:
if call.function.name != "get_weather":
raise ValueError(f"Tool not allowed: {call.function.name}")
args = json.loads(call.function.arguments)
if (
not isinstance(args, dict)
or set(args) != {"city", "unit"}
or not isinstance(args["city"], str)
or not args["city"].strip()
or args["unit"] != "celsius"
):
raise ValueError("city must be a non-empty string and unit must be celsius")
result = get_weather(args["city"], args["unit"])
messages.append({
"role": "tool",
"tool_call_id": call.id,
"content": json.dumps(result, ensure_ascii=False),
})
if not assistant.tool_calls:
raise RuntimeError("The model did not call a tool; check the model and schema")
final = client.chat.completions.create(
model="qwen2.5:3b", messages=messages, tools=TOOLS
)
print(final.choices[0].message.content)
assert assistant.tool_calls[0].function.name == "get_weather"
assert any(message["role"] == "tool" for message in messages)
python hello_tool.py
This uses the OpenAI Python SDK connected to Ollama’s compatible Chat Completions endpoint; data is not sent to the OpenAI cloud. additionalProperties: false helps with the schema, but Ollama and OpenAI strict-mode guarantees are not identical; the program must still validate.
If the model does not call the tool, keep the question, model, and schema unchanged and rerun three times, recording the success count; do not declare that the model “does not support it” after one failure.
Path B: Complete Anthropic round trip (reserve $0.05 first per run)
import json
import os
import anthropic
client = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])
tools = [{
"name": "get_weather",
"description": "Get demonstration weather data for a specified city",
"input_schema": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
"additionalProperties": False,
},
}]
def get_weather(city: str) -> dict:
return {"city": city, "temperature": 26, "unit": "celsius"}
messages = [{"role": "user", "content": "What is the temperature in Taipei now?"}]
first = client.messages.create(
model="claude-haiku-4-5-20251001",
max_tokens=512,
tools=tools,
messages=messages,
)
messages.append({"role": "assistant", "content": first.content})
tool_results = []
for block in first.content:
if block.type == "tool_use":
if block.name != "get_weather":
raise ValueError(f"Tool not allowed: {block.name}")
if (
set(block.input) != {"city"}
or not isinstance(block.input["city"], str)
or not block.input["city"].strip()
):
raise ValueError("get_weather requires one string city")
result = get_weather(block.input["city"])
tool_results.append({
"type": "tool_result",
"tool_use_id": block.id,
"content": json.dumps(result, ensure_ascii=False),
})
if not tool_results:
raise RuntimeError(f"No tool request; stop_reason={first.stop_reason}")
messages.append({"role": "user", "content": tool_results})
final = client.messages.create(
model="claude-haiku-4-5-20251001",
max_tokens=512,
tools=tools,
messages=messages,
)
print("\n".join(block.text for block in final.content if block.type == "text"))
An Anthropic client-tool failure must use the corresponding tool_use_id and add "is_error": true. Do not insert a tool result into the system prompt.
Exercise 2: Multi-Tool Selection¶
After finishing, the model chooses one of calculator and get_weather, and the program dispatches only names in the allowlist.
First action: Run this mock test directly; no key is needed:
python examples/stage-3/02-multi-tool-selection/test.py
Expand Path A/Path B, observation points, and budget
- Path A README (Ollama): run
python starter.py. - The same folder’s
starter_anthropic.pyis Path B; runpython test_anthropic.pyto validate the message shape with a mock first. - Observe
tool_calls[0].function.name, then confirm that the program rejects unknown names. - Do not dispatch tools with
globals()[model_name]()oreval().
Path A has $0 in API cost; reserve $0.05 for one Path B round first.
Structured Output (Structured Outputs / JSON mode) ⭐ function calling’s twin¶
Function Calling means “ask a program to do something”; Structured Output means “ask the model to put data into a fixed shape.” Both use schemas, but their purposes differ.
Expand strict mode, JSON mode, and common limitations
- JSON mode usually guarantees only that the response can be parsed as JSON; it does not necessarily follow your field rules.
- Structured Output constrains the schema shape within provider-supported limits; refusal, truncation, or semantic errors can still occur.
- OpenAI strict mode requires every object to set
additionalProperties: falseand list all properties as required; Chat Completions is still not strict by default. - Anthropic strict tool use has different schema and message formats from OpenAI; do not copy flag names directly.
- Ollama/other compatible endpoints vary by model and version. Validate with a fixed eval; do not infer identical behavior from “compatible.”
For Python model-based schema management, see 567-labs/instructor; for constrained decoding, see dottxt-ai/outlines. Whichever you use, the program must handle parsing and semantic errors.
Exercise 3: Implement ReAct from Scratch (No Framework)¶
After finishing, you will have a minimal Agent Loop: the model can call tools multiple times, but it always stops after the limit.
First action: Run the test that needs no key:
python examples/stage-3/03-react-from-scratch/test.py
Expand the 13-line loop, two paths, and completion conditions
for step in range(MAX_STEPS):
response = ask_model(messages, tools)
calls = read_tool_calls(response)
if not calls:
return read_final_text(response)
for call in calls:
name, args, call_id = validate_call(call)
result = TOOL_IMPL[name](**args)
messages.append(make_tool_result(call_id, result))
raise RuntimeError(f"Agent exceeded {MAX_STEPS} steps and stopped")
The real program must also put the assistant’s Tool Call back into history and handle refusal, max tokens, timeout, unknown tools, JSON parsing, and tool exceptions. The complete two paths are in 03-react-from-scratch.
Record a trace as action / observation / final or a short verifiable summary; do not make private Chain-of-Thought a logging contract.
Path A has $0 in API cost; reserve $0.05 for one Path B loop first. Completion condition: tests prove that “no tool call stops” and “exceeding MAX_STEPS raises an error.”
Exercise 4: Multi-Step Reasoning Task¶
After finishing, the same loop first checks data and then calculates, with a corresponding call ID and result for every step.
First action: Copy the test command:
python examples/stage-3/04-multi-step-reasoning/test.py
Expand the task, comparison method, and budget
Example task: “Check Taipei’s temperature, then convert it to Fahrenheit.” Use separate get_weather and celsius_to_fahrenheit tools. Do not secretly combine the two steps into one fake tool; this exercise observes whether the model continues from the previous result.
The complete two paths are in 04-multi-step-reasoning. When comparing models, keep the prompt, tools, schema, MAX_STEPS, and test cases fixed; rerun at least five times and record success rates and failure types.
Path A has $0 in API cost; reserve $0.10 for multiple Path B requests. A larger model may be more stable, or merely more expensive; use an eval to decide.
Exercise 5: Error Handling¶
After finishing, the program returns tool errors that the model can correct, while clearly stopping on transport, parsing, or limit errors.
First action: Run both mock tests:
python examples/stage-3/05-error-handling/test.py
python examples/stage-3/05-error-handling/test_anthropic.py
Expand error categories, bounded retry, and budget
| Error | What the program does first | Send back to the model? |
|---|---|---|
| Network timeout/rate limit | Retry with a bound; record the error | Usually not at first |
| Tool Call JSON parsing failure | Do not execute the tool; report a format error | Yes, as an error result |
| Unknown tool/unauthorized argument | Reject execution; leave an audit log | Yes, but never relax permissions |
| Tool cannot find data | Return a clear, minimal semantic error | Yes, so the model can revise or give up |
MAX_STEPS/cost limit reached |
Stop immediately | Do not retry |
Anthropic’s failed tool_result uses "is_error": true. On the OpenAI-compatible path, structured errors can go in the role: tool content, but the application must still limit retries.
The complete two paths are in 05-error-handling. Path A has $0 in API cost; reserve $0.10 for one Path B error-recovery round.
Exercise 6: Function Schema Design (Fixing a Bad Schema)¶
After finishing, you will compare two schemas with the same set of questions and identify improvements to descriptions, fields, enums, or constraints.
First action: Run the bad and good mock tests directly:
python examples/stage-3/06-schema-design/test.py
python examples/stage-3/06-schema-design/test_anthropic.py
Expand the five rules, eval card, and budget
- Use a clear verb plus noun for a tool name, such as
get_weather. - Say when to use the tool and when not to use it in the description.
- Give every field a clear name, type, and example.
- Use
enum, ranges, andadditionalProperties: falseto constrain inputs explicitly when possible. - The schema owns only the interface; the program still validates permissions, business rules, and data safety.
The complete two paths are in 06-schema-design, with a quick reference in resources/schema-design-cheatsheet.en.md.
Copy this result card directly; you do not need to draw a blank table first:
Fixed prompt: ________________
Bad schema | success __ / 5 | main error: ________________
Good schema | success __ / 5 | main improvement: ________________
Conclusion | most helpful field: ________________
Do not write “a certain model almost always fails.” Path A has $0 in API cost; reserve $0.25 for five Path B comparisons first.
🎒 Recommended Mini-Project: A Safe Weather Helper¶
Connect Exercises 1–6, keeping only two read-only tools: get_weather and convert_temperature. Add an allowlist, argument validation, MAX_STEPS, a timeout, error results, and a five-question eval.
The minimum deliverable is agent.py, test_agent.py, eval_cases.json, and one result card. Get the mock tests passing before running a local model; do not start with payment, file-deletion, or email tools.
🪞 Reflection (Reflexion / Self-Refine) — Concept + Routing¶
Expand the relationship between Reflection, Reflexion, Self-Refine, and memory
- Reflection is the broad term: inspect the previous round and improve the next one.
- Reflexion often writes failures, feedback, and the next strategy into reusable text records.
- Self-Refine often improves one output through a “generate → critique → rewrite” cycle.
- These are sibling patterns to ReAct; they are not Tool Use and do not necessarily need persistent memory.
This chapter covers a single-session loop only. For carrying failed experiences across sessions, go to Stage 06 Reflection Memory; for fuller planning, verification, and long-running execution, go to Stage 07.5.
🎯 Curated Projects¶
Complete one five-star route first: official docs → Exercises 1–3 → one from-scratch implementation. The full table is a toolbox, not a list of 21 tasks.
Resources checked: 2026-08-27 UTC
Ratings indicate this Stage’s learning priority, not popularity:
⭐⭐⭐⭐⭐= skipping it will block this chapter’s route;⭐⭐⭐⭐= recommended early;⭐⭐⭐= read if needed;⭐⭐= historical or niche context.
| Category | Resource | Do first | Status / license | Rating |
|---|---|---|---|---|
| Official core docs | Anthropic — How Tool Use Works | Start with the five-step client-tool round trip. | Official docs | ⭐⭐⭐⭐⭐ |
| Anthropic — Handle Tool Calls | Look at call IDs, results, and is_error. | Official docs | ⭐⭐⭐⭐⭐ | |
| Ollama — Tool Calling | Run the single-tool and agent-loop examples once. | Official docs | ⭐⭐⭐⭐⭐ | |
| OpenAI — Function Calling | Compare function schemas and strict mode. | Official docs | ⭐⭐⭐⭐ | |
| Google Gemini — Function Calling | Compare sequential/parallel calls when you need Gemini. | Official docs | ⭐⭐⭐ | |
| ReAct paper | Read the abstract and method diagram first. | Original paper; arXiv | ⭐⭐⭐⭐ | |
| Official courses and examples | Anthropic Courses — Tool Use | Read the older Tool Use notebook; use the Cookbook in the next row when building. | Archived official course; upstream provides no SPDX | ⭐⭐⭐⭐ |
| Anthropic Tool Use Cookbook | Move from one tool to parallel tools. | Maintained; MIT | ⭐⭐⭐⭐⭐ | |
| Anthropic Quickstarts | After the exercises, see how a full app connects tools. | Maintained; MIT | ⭐⭐⭐⭐ | |
| Microsoft AI Agents for Beginners | Choose a chapter if you want another complete course. | Maintained; MIT | ⭐⭐⭐⭐ | |
| From-scratch implementations | pguso/ai-agents-from-scratch | Use Ollama to compare with Exercise 3’s loop. | Maintained; MIT | ⭐⭐⭐⭐⭐ |
| arunpshankar/react-from-scratch | Read later for Gemini/Reflection variants. | Updates slowed (last push 2025-05); Apache-2.0 | ⭐⭐⭐ | |
| mattambrogi/agent-implementation | Use only to read through a minimal teaching toy line by line. | Historical reference (last push 2024-01); upstream provides no SPDX | ⭐⭐ | |
| lsdefine/GenericAgent | Compare it later if you want to see a small framework. | Maintained; MIT | ⭐⭐⭐ | |
| Framework / CodeAct comparisons | Hugging Face Smolagents | Compare CodeAct after completing the JSON-tool loop. | Maintained; Apache-2.0 | ⭐⭐⭐⭐ |
| QuantaLogic | Read later when you need a second CodeAct implementation. | Updates slower (last push 2025-12); Apache-2.0 | ⭐⭐⭐ | |
| LangChain ReAct Agent | See how a framework wraps the loop you wrote yourself. | Maintained; MIT | ⭐⭐⭐ | |
| Chinese chapter-style textbooks | datawhalechina/hello-agents | Use this route for complete Chinese chapters. | Maintained; upstream metadata provides no SPDX | ⭐⭐⭐⭐⭐ |
| jjyaoao/HelloAgents | Run the code alongside the textbook; check the matching branch first. | Maintained; upstream metadata provides no SPDX | ⭐⭐⭐⭐⭐ | |
| Structured Output tools | 567-labs/instructor | Read it for typed models, validation, and retry. | Former jxnl/instructor redirects here; MIT | ⭐⭐⭐⭐ |
| dottxt-ai/outlines | Read it to study constrained decoding locally. | Maintained; Apache-2.0 | ⭐⭐⭐⭐ |
✅ Self-Check Before Stage 4¶
- I can explain
schema → call → execute → result → answerin my own words. - I can distinguish Tool Call, Tool Result, and Structured Output.
- My program dispatches only allowlisted tools, validates arguments, and has
MAX_STEPS. - I ran Exercises 1–3 and saw at least one successful and one error path.
- When comparing models or schemas, I used the same test set and explicit scores.
Once these are done, enter Stage 4 — Workflow Graphs & Agent Frameworks. If you still cannot explain the full round trip, rerun Exercise 1; you do not need to reread the whole chapter.