The Agent Loop
Perceive, think, act, observe - repeated: the loop at the heart of every agent, ReAct, and how and when it must stop.
What is it?
The agent loop is the small piece of code that turns tool use into an agent. It is often described as four phases repeated until done:
- Perceive: the model receives the goal plus everything that has happened so far (the full message history).
- Think: it reasons about what to do next. Current Claude models do this with built-in adaptive thinking; older systems prompted the model to write its reasoning out.
- Act: it requests one or more tool calls (or decides it is finished).
- Observe: your code runs the tools and appends their results; the next iteration starts with this new information.
This idea was popularised by ReAct (Reason + Act), a prompting technique where the model alternated lines of 'Thought:', 'Action:' and 'Observation:' in plain text and the program parsed the Action lines. Modern APIs replace the text parsing with native tool_use blocks, but the rhythm - reason, act, observe, repeat - is the same.
The loop in code is short: call the model; append its reply to the history; if it asked for tools, run them, append the results as one user message, and loop; otherwise, stop. The interesting part is when to stop, and the API tells you through the response's stop_reason:
end_turn: the model finished its answer. Normal exit.tool_use: the model wants tools run. Keep looping.max_tokens: the reply was cut off because it hit yourmax_tokenslimit. The answer is incomplete - raise the limit or ask the model to continue; do not treat it as a final answer.stop_sequence: generation hit one of your custom stop strings.refusal: the model declined for safety reasons. Stop and surface this; do not retry the same request in a loop.pause_turn: used with server-side tools that run on the provider's side; the model paused a long turn. Send the conversation back (with the paused assistant content appended) to let it continue.
On top of the model's own stopping, you must add stop conditions, because a model can loop - repeating the same search, retrying a broken tool, or wandering:
- Max steps (iterations): the single most important guard. 10-30 is typical for interactive agents.
- Budget: stop when total tokens or cost passes a limit (sum
usage.input_tokensandusage.output_tokenseach step). - Wall-clock timeout: stop after N seconds.
- Repetition detection: stop (or warn the model) when the same tool is called with the same input several times.
- Human checkpoints: pause for approval before risky actions, or when the agent is stuck.
When a guard triggers, end gracefully: return what was learned so far and say why it stopped, rather than crashing. Log every step - the model's text, each tool name and input, each result (truncated), stop_reason and token usage - so you can replay and debug runs later.
Explain like I'm 10
A detective solving a case: look at the evidence (perceive), form a theory (think), go interview a witness or check a record (act), note what was learned (observe), and repeat. A good detective also has a boss who says 'you have three days and a fixed budget - report back either way'. That boss is your max-steps and budget guard.
Examples
Classic text ReAct: parse Thought / Action / Observation (runnable)
// Before native tool calling, ReAct agents wrote actions as text and the program parsed them.
const kb = { "capital of france": "Paris", "population of paris": "about 2.1 million (city proper)" };
const tools = {
search: (q) => kb[q.toLowerCase()] || "no result",
};
// Scripted model output for each turn (a real LLM would generate these lines).
const modelTurns = [
"Thought: I need the capital first.\nAction: search[capital of France]",
"Thought: Now I need its population.\nAction: search[population of Paris]",
"Thought: I have both facts.\nFinal Answer: The capital is Paris, with about 2.1 million people in the city proper.",
];
let transcript = "Question: What is the capital of France and how many people live there?";
for (let step = 0; step < 5; step++) {
const output = modelTurns[step];
transcript += "\n" + output;
const lines = output.split("\n");
const final = lines.find((l) => l.startsWith("Final Answer:"));
if (final) { console.log(transcript); console.log("--> stopped: model gave a final answer at step", step + 1); break; }
const action = lines.find((l) => l.startsWith("Action:"));
const name = action.slice("Action: ".length, action.indexOf("["));
const arg = action.slice(action.indexOf("[") + 1, action.lastIndexOf("]"));
transcript += "\nObservation: " + tools[name](arg);
}The loop is the same as today's: model output -> run action -> append observation -> repeat. Native tool_use blocks simply replace the fragile string parsing with structured JSON.
Guards in action: a model stuck in a loop (runnable)
// A buggy "model" that keeps searching for the same thing forever.
function stuckModel(history) {
return { stop_reason: "tool_use", call: { name: "search", input: { q: "refund policy" } }, tokens: 1200 };
}
function runLoop(model, { maxSteps, maxTokens, maxRepeats }) {
const history = [];
const seen = {};
let tokensUsed = 0;
for (let step = 1; step <= maxSteps; step++) {
const resp = model(history);
tokensUsed += resp.tokens;
if (resp.stop_reason === "end_turn") return "done: " + resp.text;
if (resp.stop_reason === "max_tokens") return "incomplete: reply was truncated";
if (resp.stop_reason === "refusal") return "refused: surfacing to the user";
if (tokensUsed > maxTokens) return "stopped: token budget exceeded at step " + step + " (" + tokensUsed + " tokens)";
const key = resp.call.name + JSON.stringify(resp.call.input);
seen[key] = (seen[key] || 0) + 1;
if (seen[key] > maxRepeats) return "stopped: repeated " + key + " " + seen[key] + " times at step " + step;
console.log("step", step, "->", resp.call.name, JSON.stringify(resp.call.input));
history.push({ call: resp.call, result: "Refunds within 30 days." });
}
return "stopped: hit max steps (" + maxSteps + ")";
}
console.log(runLoop(stuckModel, { maxSteps: 10, maxTokens: 100000, maxRepeats: 2 }));
console.log(runLoop(stuckModel, { maxSteps: 3, maxTokens: 100000, maxRepeats: 99 }));
console.log(runLoop(stuckModel, { maxSteps: 50, maxTokens: 5000, maxRepeats: 99 }));Three different guards catch the same runaway model: repetition detection, max steps, and a token budget. Production loops usually combine all three.
A production-shaped loop handling every stop_reason
import anthropic, json, logging
logging.basicConfig(level=logging.INFO)
log = logging.getLogger("agent")
client = anthropic.Anthropic()
def agent_loop(task, tools, run_tool, system, max_steps=15, max_total_tokens=400_000):
messages = [{"role": "user", "content": task}]
total = 0
for step in range(1, max_steps + 1):
resp = client.messages.create(model="claude-opus-5-5", max_tokens=16000,
system=system, tools=tools, messages=messages)
total += resp.usage.input_tokens + resp.usage.output_tokens
log.info("step=%d stop=%s tokens_so_far=%d", step, resp.stop_reason, total)
messages.append({"role": "assistant", "content": resp.content})
if resp.stop_reason == "end_turn":
return "".join(b.text for b in resp.content if b.type == "text")
if resp.stop_reason == "max_tokens":
return "[incomplete] the reply hit max_tokens; raise the limit or split the task"
if resp.stop_reason == "refusal":
return "[refused] the model declined this request"
if resp.stop_reason == "pause_turn":
continue # server tool paused: send history back so it resumes
if resp.stop_reason != "tool_use":
return f"[stopped] unexpected stop_reason {resp.stop_reason}"
results = []
for block in resp.content:
if block.type != "tool_use":
continue
log.info(" call %s %s", block.name, json.dumps(block.input)[:200])
try:
out = run_tool(block.name, block.input)
results.append({"type": "tool_result", "tool_use_id": block.id, "content": out})
except Exception as e:
results.append({"type": "tool_result", "tool_use_id": block.id,
"content": f"Error: {e}", "is_error": True})
messages.append({"role": "user", "content": results})
if total > max_total_tokens:
return "[stopped] token budget exceeded"
return f"[stopped] no answer after {max_steps} steps"Each exit path is explicit: normal finish, truncated, refused, paused, unknown, budget, and max steps. The log line per step is your future debugging lifeline.
How it works
Every iteration re-sends the complete history, so step N's request contains the original task, every earlier assistant message (with its tool_use blocks), and every tool result. This is how the model 'remembers' what it has already tried. It is also why long runs get slower and more expensive, and why large tool outputs should be trimmed.
The history must stay well-formed: roles alternate, every tool_use is followed in the very next user message by a tool_result with the same id, and tool results come first in that user message. Breaking these rules produces API errors - a frequent bug when people prune or edit history.
Reliability compounds. If each step is right 95% of the time, a 10-step run with no recovery succeeds only about 60% of the time (0.95 to the power 10). Agents survive this by observing results and correcting course - which is why clear error messages from tools and good feedback signals (tests, validators) matter so much.
+-----------------------------+
| messages = [task] |
+--------------+--------------+
v
+---------------------+
+------->| call model |
| +----------+----------+
| v
| stop_reason?
| tool_use | end_turn / max_tokens /
| | refusal / guard hit
| v \
| run tools, append ONE --> return answer
| user msg of tool_results or reason
| |
+-- step < max --+Why does it exist?
A single model call can only use what is already in the prompt. Real tasks need several rounds of gathering information and acting on it. The loop lets the model use each observation to decide its next move, which is what makes multi-step problem solving possible. Explicit stop conditions exist because models do not always know when to give up.
When to use it
Use a loop whenever the model may need more than one round of tools to finish - research, debugging, data exploration, multi-step support actions. Write it yourself at least once to understand it; later, an SDK helper such as a tool runner can manage it for you.
When not to use it
If you know the model will always need exactly one tool call (or a fixed sequence), skip the loop and write a workflow - two explicit calls are easier to test. If latency matters more than flexibility, cap the loop at 1-3 steps or avoid it.
Common mistakes
No max-steps limit - the single most expensive agent bug.
Treating stop_reason max_tokens as a finished answer, shipping a truncated reply.
Retrying in a loop after a refusal instead of stopping and explaining.
Only checking
stop_reason == 'end_turn'and silently exiting on anything else, hiding problems.Not appending the assistant's full content before the tool results, breaking the tool_use/tool_result pairing.
Logging nothing, or logging only the final answer, so failed runs cannot be diagnosed.
Letting tool results grow without limit so each step re-sends megabytes of text.
Practice exercises
- Easy:
Compute the success rate of a 5-step, 10-step and 20-step run when each step is independently correct 90% of the time. What does this suggest about long agent runs?
- Easy:
Add a fourth guard to the runnable guard demo: a wall-clock limit using Date.now(). Test it by making the fake model slow with a busy loop.
- Medium:
Extend the ReAct demo with a
calculatetool and a question that needs search then arithmetic. Keep the text format and parsing. - Medium:
Modify the Python loop so that after 3 identical tool calls it injects a user note: 'You have called X with the same input 3 times; try a different approach or answer with what you have.'
- Hard:
Make the Python loop write a JSON-lines trace file with one record per step (step, stop_reason, tool calls, truncated results, tokens, elapsed ms). Write a 20-line script that replays a trace as a readable timeline.
Interview questions
Describe the agent loop.
Call the model with the full history and tool definitions; append its response; if stop_reason is tool_use, execute each requested tool, append one user message with all tool_results, and repeat; otherwise stop. Wrap it with guards: max steps, token or cost budget, timeouts, repetition detection, and approval for risky actions.
What is ReAct?
Reasoning and Acting: a prompting pattern where the model interleaves reasoning ('Thought'), actions ('Action' calling a tool), and the results ('Observation') in a loop. It showed that letting models act and observe improves multi-step tasks. Native tool calling is the structured modern form of the same idea.
What stop reasons should an agent loop handle?
end_turn (done), tool_use (run tools and continue), max_tokens (truncated - not a valid final answer), stop_sequence (custom stop string), refusal (stop and surface), and pause_turn (server tools paused; resend to continue). Anything unexpected should be logged and handled explicitly.
Why do long agent runs fail more often?
Errors compound: each step has some chance of a wrong action, and later steps build on earlier ones. The context also grows, adding noise and cost. Mitigations are feedback signals, clear tool errors, smaller tasks, step limits and checkpoints.
Why does every step re-send the whole conversation?
The API is stateless. The model only knows what is in the current request, so to remember earlier tool calls and results you must include them each time. Prompt caching can reduce the cost of the repeated prefix.
What should you log per step?
Step number, stop_reason, model text, each tool name and input, each result (truncated) and whether it errored, token usage and latency. This lets you reconstruct the trajectory and find where a run went wrong.