What is an AI Agent?

An LLM that decides its own next steps in a loop, using tools, until a goal is done - and when a simpler design is better.

What is it?

An AI agent is a program where a large language model (LLM) is given a goal, a set of tools (functions it may ask your code to run, such as 'read a file' or 'search the docs'), and a loop. On every turn the model looks at everything that has happened so far and decides what to do next: call a tool, call several tools, or stop and give an answer. Your code runs the tools the model asked for, sends the results back, and the model decides again.

The key word is decides. In an agent, the model chooses the sequence of steps at run time. You do not know in advance whether it will take two steps or twelve, or which tools it will pick. That flexibility is the whole point - and also the whole risk.

It helps to place agents on a ladder of three designs that people often confuse:

  • Chatbot (single call): text in, text out. One request to the model, one reply. No tools, no loop. Example: 'rewrite this email to sound friendlier'.
  • Workflow: an LLM is used inside a sequence of steps that your code defines. The path is fixed or chosen by simple rules: classify the ticket, then route it, then draft a reply, then check the draft. The model fills in the blanks; the program holds the steering wheel.
  • Agent: the model holds the steering wheel. It picks which tool to use, reads the result, and keeps going until it judges the task done (or until your safety limits stop it).

A useful definition from practitioners: workflows are systems where LLMs and tools are orchestrated through predefined code paths; agents are systems where the LLM dynamically directs its own process and tool usage. Both are useful. Neither is automatically better.

Every agent is built from the same four parts:

  • Model: the LLM that reasons and chooses actions.
  • Tools: functions with a name, a description, and an input schema (a JSON Schema describing the arguments). The model can only request a tool; your code actually runs it.
  • Loop: call the model, run any requested tools, append the results to the conversation, call the model again - until it stops asking for tools.
  • Guardrails: a maximum number of steps, timeouts, permission checks, human approval for risky actions, and logging so you can see what happened.

When NOT to build an agent is the most important lesson on this page. Start with the simplest thing that works: a single well-written prompt. If that is not enough, add retrieval (fetching relevant documents) or a fixed workflow. Only reach for an agent when the steps genuinely cannot be known in advance - because agents cost more (many model calls per task), are slower, are harder to test (each run can take a different path), and can make compounding mistakes (a wrong step early leads every later step astray).

Good agent tasks share a few traits: the task is open-ended, the number of steps varies, the model needs feedback from the environment (test results, search hits, file contents) to make progress, and mistakes are cheap to detect and recover from. Coding assistants that edit files and run tests, research assistants that search and read iteratively, and support assistants that look up orders and take bounded actions are classic examples.

Explain like I'm 10

A chatbot is a calculator: press keys, get one answer. A workflow is a recipe: step 1 chop, step 2 fry, step 3 serve - the cook follows the card. An agent is a chef you hand a goal ('make dinner for four, one guest is vegetarian') and a stocked kitchen: the chef opens the fridge, tastes, adjusts, and decides each next step. You get flexibility, but you also need house rules - a budget, no using the expensive knives without asking, and a deadline.

Examples

Chatbot vs workflow vs agent, side by side (runnable, offline)

// A fake LLM so this runs without a network. Real code would call an API.
function llm(prompt) {
  if (prompt.startsWith("Classify:")) return prompt.includes("refund") ? "billing" : "technical";
  if (prompt.startsWith("Reply:")) return "Thanks for writing in - we are on it.";
  return "Our opening hours are 9am-5pm.";
}

// 1) CHATBOT: one call, text in -> text out.
console.log("chatbot  :", llm("What are your hours?"));

// 2) WORKFLOW: YOUR code fixes the steps; the LLM fills in each one.
const ticket = "I was charged twice, I want a refund for order 42";
const team = llm("Classify: " + ticket);           // step 1 (LLM)
const draft = llm("Reply: " + ticket);              // step 2 (LLM)
console.log("workflow : routed to", team, "| draft:", draft);

// 3) AGENT: the MODEL picks the next action each turn, in a loop.
const orders = { 42: { paid: 2, amount: 30 } };
const tools = {
  lookup_order: (id) => JSON.stringify(orders[id]),
  issue_refund: (id) => { orders[id].paid -= 1; return "refunded " + orders[id].amount; },
};
// A scripted "model": it decides based on what it has observed so far.
function agentModel(history) {
  const seen = history.join(" ");
  if (!seen.includes("paid")) return { tool: "lookup_order", arg: 42 };
  if (seen.includes('"paid":2') && !seen.includes("refunded")) return { tool: "issue_refund", arg: 42 };
  return { final: "You were charged twice; I refunded one charge of 30." };
}
const history = [];
for (let step = 1; step <= 5; step++) {               // max-steps guard
  const action = agentModel(history);
  if (action.final) { console.log("agent    : FINAL ->", action.final); break; }
  const observation = tools[action.tool](action.arg);
  console.log("agent    : step", step, action.tool + "(" + action.arg + ") ->", observation);
  history.push(observation);
}

Notice who decides the order of steps. In the workflow, the program hard-codes classify-then-reply. In the agent, the scripted model looked at the order first, saw a double charge, chose to refund, then chose to stop. A real LLM makes those choices from the conversation instead of a script.

The simplest LLM call (a chatbot turn) with the Anthropic SDK

# pip install anthropic   and set ANTHROPIC_API_KEY in your environment
import anthropic

client = anthropic.Anthropic()

resp = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=16000,
    system="You are a friendly support assistant for an online shop.",
    messages=[{"role": "user", "content": "What is your returns window?"}],
)
for block in resp.content:
    if block.type == "text":
        print(block.text)

# Before building an agent, ask: is this already good enough?
# If one call (plus maybe retrieved documents) solves it, stop here.

This is the baseline every agent project should beat. An agent is this same call placed inside a loop, with tools attached.

Decision checklist: do I need an agent?

1. Can one prompt (with good instructions + examples) do it?   -> single call
2. Does it need facts from your documents?                      -> add retrieval (RAG)
3. Are the steps always the same, in the same order?            -> workflow (chain)
4. Are there a few known kinds of input needing different handling? -> workflow (routing)
5. Are the steps unknown until you see intermediate results,
   and can the model check its own progress (tests, search hits)? -> agent
6. Are mistakes expensive or irreversible (money, deletes, emails)?
   -> agent only with human approval + tight permissions, or don't

Walk down the list and stop at the first 'yes'. Most production LLM features stop at step 1-4.

How it works

Under the hood an agent is an ordinary program. The model never touches your files or databases directly. It produces a structured request such as {name: "read_file", input: {path: "todo.txt"}}. Your code checks that request, runs the real function, and appends the output to the conversation as a tool result. Then it calls the model again with the whole conversation, because LLM APIs are stateless - the model remembers nothing between calls unless you send it.

Each pass through the loop is called a step or turn. The loop ends when the model replies without asking for a tool (it considers the task done), when a safety limit is reached (maximum steps, time, or cost), or when something needs a human decision.

Because every step re-sends the growing conversation, cost and latency grow with the number of steps. A 10-step agent run can easily cost 10-30 times a single call. That is why the advice is always: prove you need the flexibility before paying for it.

  CHATBOT          WORKFLOW                AGENT
  -------          --------                -----
  input            input                   goal
    |                |                       |
  [LLM]           [LLM step 1]         +-> [LLM decides] --+
    |                |                 |       |           |
  output          [LLM step 2]         |   tool call?    answer
                     |                 |       |
                  [check]              |   [run tool]
                     |                 |       |
                  output               +-- result
  1 call          fixed path           model picks path

Why does it exist?

Many real tasks cannot be scripted in advance. Fixing a bug means reading code, forming a guess, running tests, and reacting to what fails. Researching a question means searching, reading, and searching again based on what you learned. A fixed pipeline cannot anticipate every branch; an LLM that can choose actions and observe results can.

Agents exist to turn a model that only talks into a system that can act - safely, within limits you define.

When to use it

Use an agent when the task is open-ended, the number and order of steps depend on intermediate results, the environment gives clear feedback (a test passes, a search returns hits), and you can bound the damage of a wrong action. Coding assistants, research assistants, data-exploration helpers and support bots with a few safe actions are typical fits.

When not to use it

Do not build an agent when a single prompt, retrieval, or a fixed workflow would do. Do not build one when every step is known in advance (use a workflow - it is cheaper, faster and testable), when latency must be very low, when the cost per request must be predictable, or when actions are irreversible and you cannot add human approval. 'Agent' is not a goal; a working, reliable feature is.

Common mistakes

  • Starting with an agent framework before trying a single good prompt - most problems never need a loop.

  • Calling any LLM feature an 'agent'; if your code fixes the steps it is a workflow, which is usually a good thing.

  • Running the loop with no maximum step count, so a confused model can burn money forever.

  • Giving the agent powerful tools (delete, send email, pay) with no permission checks or human approval.

  • Expecting the model to 'remember' earlier calls automatically - the API is stateless, you must send the history.

  • Not logging each step, which makes it impossible to understand why an agent did something strange.

  • Assuming more autonomy means better results; reliability usually drops as the number of free choices grows.

Practice exercises

  1. Easy:

    For each of these, say whether a single call, a workflow, or an agent fits best, and why: translating a paragraph; triaging support tickets into 4 queues; fixing a failing unit test in an unknown codebase; writing a weekly report from a fixed set of 3 dashboards.

  2. Easy:

    In the runnable demo, change the order data so the customer was charged only once. Predict what the scripted agent does, then run it and check.

  3. Medium:

    Extend the demo's agent with a third tool, send_email, and make the loop refuse to run it unless a variable approved is true. Print what the agent does in both cases.

  4. Medium:

    Write down the guardrails (limits, permissions, approvals, logs) you would require before letting an agent issue refunds in a real shop.

  5. Hard:

    Pick a task from your own work. Build it first as a single prompt, then as a 2-3 step workflow. Write a short note on whether an agent would actually add value, with evidence from the failures you saw.

Interview questions

What is the difference between a workflow and an agent?

In a workflow, your code defines the sequence of LLM calls and tool calls (fixed steps or rule-based branching). In an agent, the LLM decides at run time which tools to call and when to stop, inside a loop. Workflows are more predictable and cheaper; agents are more flexible for open-ended tasks.

What are the core components of an agent?

A model that reasons and chooses actions, a set of tools with names, descriptions and input schemas, a loop that runs requested tools and feeds results back, and guardrails such as maximum steps, permissions, human approval and logging.

When would you advise against building an agent?

When a single prompt, retrieval, or a fixed workflow solves the task; when steps are known in advance; when latency or cost must be low and predictable; or when actions are irreversible and cannot be gated by a human. Agents add cost, latency, nondeterminism and compounding errors.

Does the model execute tools itself?

No. The model only emits a structured request naming a tool and its arguments. Your application validates it, executes the real function with its own permissions, and sends the result back as a tool result message.

Why do agents cost more than a single call?

Each step is another model call, and because the API is stateless each call re-sends the growing conversation, including all previous tool calls and results. Token usage grows with the number of steps and with the size of tool outputs.

What makes a task a good fit for an agent?

Open-ended goals where the steps depend on intermediate results, clear feedback signals from the environment (tests, search results, errors), the ability to recover from mistakes, and actions whose risk can be bounded with permissions or approval.