Evaluating Agents
Test agents on task success, trajectories, tool-call accuracy and step/cost budgets, with repeated trials and regression suites.
What is it?
Evaluating an agent is harder than evaluating a single LLM call because an agent takes many steps, chooses its own tools, and can reach a correct answer by a bad route (or a wrong answer by a reasonable one). You therefore grade at three levels: the outcome, the trajectory, and the budget.
- Task success (outcome): did the agent achieve the goal? Best checked against the state of the world after the run (the file was created, the ticket was closed, the database row has the right value), or with a deterministic check on the final answer. Use an LLM judge only when the outcome is genuinely free text.
- Trajectory: the ordered list of steps the agent took: each model turn, each tool call with its arguments, each tool result. Trajectory checks ask: did it call the tools it needed? Did it avoid forbidden tools? Did it ask for approval before a destructive action? Did it loop?
- Tool-call accuracy: for each call, was the right tool chosen and were the arguments valid and correct (right order id, right date format)? You can score this like retrieval: precision (how many calls were necessary and correct) and recall (how many required calls were made).
- Budgets: number of steps, total tokens (and therefore cost), and wall-clock time. An agent that succeeds in 40 steps when 4 would do is a production problem.
Agents are non-deterministic: the same task can succeed on one run and fail on the next. So you run each task several times and report a success rate, not a single pass/fail. Two useful ways to summarise repeated trials: pass@k (did at least one of k attempts succeed? useful when a human can pick the best) and pass^k (did all k attempts succeed? the right bar for an unattended agent that must be reliable every time).
Test tasks should run in a sandbox with fake or recorded tools: a mock order API, a temporary directory, a seeded test database. That makes runs safe (no real emails sent), repeatable, and lets you check the end state precisely.
As with RAG, the eval suite becomes a regression suite: every bug you find in production becomes a new test task, and every change to the prompt, tools or model is run against the whole suite before it ships.
Explain like I'm 10
Evaluating an agent is like a driving test, not a written exam. The examiner checks that you arrived (task success), but also how you drove: did you check mirrors, signal, stop at red lights (trajectory), and did you take a sensible route instead of circling the block ten times (budget). One clean drive is not proof you are safe; a good examiner wants consistent behaviour.
Examples
Grading recorded trajectories: outcome, tools, budgets (runnable)
// Test cases describe what a good run must satisfy.
const cases = [
{ id: "order-status", task: "What is the status of order 1042?",
expect: { answerIncludes: "shipped", mustCall: [{ name: "get_order", args: { id: "1042" } }],
mustNotCall: ["cancel_order"], maxSteps: 4, maxTokens: 4000 } },
{ id: "cancel-order", task: "Cancel order 2001 please.",
expect: { answerIncludes: "cancelled", mustCall: [{ name: "request_approval" }, { name: "cancel_order", args: { id: "2001" } }],
mustNotCall: [], maxSteps: 5, maxTokens: 5000 } },
];
// Recorded runs (in real life: captured by your agent harness).
const runs = {
"order-status": { final: "Order 1042 has shipped and arrives Friday.", steps: [
{ tool: "search_docs", args: { q: "order status" }, tokens: 900 },
{ tool: "get_order", args: { id: "1042" }, tokens: 1100 },
{ tool: null, tokens: 700 } ] },
"cancel-order": { final: "Done, order 2001 is cancelled.", steps: [
{ tool: "cancel_order", args: { id: "2001" }, tokens: 1200 }, // skipped approval!
{ tool: null, tokens: 600 } ] },
};
const argsMatch = (want, got) => Object.keys(want || {}).every(k => String(got[k]) === String(want[k]));
function grade(c, run) {
const calls = run.steps.filter(s => s.tool);
const checks = {};
checks.success = run.final.toLowerCase().includes(c.expect.answerIncludes);
// Required calls must appear in order (a subsequence of the actual calls).
let i = 0;
for (const call of calls) {
const want = c.expect.mustCall[i];
if (want && call.tool === want.name && argsMatch(want.args, call.args)) i++;
}
checks.requiredToolsInOrder = i === c.expect.mustCall.length;
checks.noForbiddenTools = !calls.some(s => c.expect.mustNotCall.includes(s.tool));
const needed = calls.filter(s => c.expect.mustCall.some(m => m.name === s.tool)).length;
checks.toolPrecision = calls.length ? +(needed / calls.length).toFixed(2) : 1;
const tokens = run.steps.reduce((a, s) => a + s.tokens, 0);
checks.withinSteps = run.steps.length <= c.expect.maxSteps;
checks.withinTokens = tokens <= c.expect.maxTokens;
const pass = checks.success && checks.requiredToolsInOrder && checks.noForbiddenTools &&
checks.withinSteps && checks.withinTokens;
return { pass, tokens, checks };
}
for (const c of cases) {
const r = grade(c, runs[c.id]);
console.log(c.id, r.pass ? "PASS" : "FAIL", "tokens=" + r.tokens, JSON.stringify(r.checks));
}The cancel-order run 'succeeded' (the answer says cancelled) but skipped the approval step, so the trajectory check fails it. Outcome-only grading would have passed a dangerous agent. Tool precision (0.5 for order-status) shows an unnecessary search call: not a failure, but a cost signal worth tracking.
Repeated trials: success rate, pass@k and pass^k (runnable)
// Deterministic pseudo-random generator so the demo prints the same thing every time.
function mulberry32(seed) {
return function () {
seed |= 0; seed = (seed + 0x6D2B79F5) | 0;
let t = Math.imul(seed ^ (seed >>> 15), 1 | seed);
t = (t + Math.imul(t ^ (t >>> 7), 61 | t)) ^ t;
return ((t ^ (t >>> 14)) >>> 0) / 4294967296;
};
}
const rand = mulberry32(42);
// Pretend each task has a hidden true reliability; a "run" succeeds with that probability.
const tasks = { "order-status": 0.95, "cancel-order": 0.8, "multi-step-refund": 0.55 };
const TRIALS = 20;
for (const [task, p] of Object.entries(tasks)) {
let ok = 0;
for (let t = 0; t < TRIALS; t++) if (rand() < p) ok++;
const rate = ok / TRIALS;
// With an estimated per-run success rate r:
const passAt3 = 1 - Math.pow(1 - rate, 3); // at least 1 of 3 succeeds
const passHat3 = Math.pow(rate, 3); // all 3 succeed
console.log(task.padEnd(18), "success", (rate * 100).toFixed(0) + "%",
" pass@3", passAt3.toFixed(2), " pass^3", passHat3.toFixed(2));
}A task that succeeds most of the time looks fine on pass@3, but pass^3 shows how often it would fail at least once across three unattended runs. These formulas assume independent runs; they are estimates, so use enough trials (tens, not two) before drawing conclusions.
Recording a trajectory from the real agent loop (Python)
import time
import anthropic
client = anthropic.Anthropic()
def run_agent(task: str, tools: list, run_tool, max_steps: int = 10) -> dict:
"""Run the manual agent loop and record everything an evaluator needs."""
messages = [{"role": "user", "content": task}]
trajectory, in_tok, out_tok = [], 0, 0
start = time.time()
resp = None
for step in range(max_steps):
resp = client.messages.create(
model="claude-opus-5-5", max_tokens=16000, tools=tools, messages=messages,
)
in_tok += resp.usage.input_tokens
out_tok += resp.usage.output_tokens
messages.append({"role": "assistant", "content": resp.content})
if resp.stop_reason != "tool_use":
break
results = []
for block in resp.content:
if block.type == "tool_use":
try:
output, is_error = run_tool(block.name, block.input), False
except Exception as e:
output, is_error = f"Error: {e}", True
trajectory.append({"step": step, "tool": block.name,
"args": block.input, "is_error": is_error})
results.append({"type": "tool_result", "tool_use_id": block.id,
"content": str(output), "is_error": is_error})
messages.append({"role": "user", "content": results})
final = "".join(b.text for b in resp.content if b.type == "text")
return {
"final": final,
"trajectory": trajectory,
"steps": len([m for m in messages if m["role"] == "assistant"]),
"input_tokens": in_tok,
"output_tokens": out_tok,
"seconds": round(time.time() - start, 2),
"stop_reason": resp.stop_reason, # "tool_use" here means we hit max_steps
}The harness wraps the same loop you built earlier and records each tool call, token usage summed across every turn, step count and stop reason. Feed this record into a grader like the runnable demo. Point run_tool at sandboxed fakes during evaluation, never at production systems.
Agent regression suite with state checks (Python, pytest)
# test_agent.py
import pytest
from harness import run_agent
from sandbox import make_sandbox # your code: fresh fake order DB + tool implementations
CASES = [
{"id": "order-status", "task": "What is the status of order 1042?",
"must_call": ["get_order"], "forbid": ["cancel_order"], "max_steps": 4},
{"id": "cancel-order", "task": "Cancel order 2001.",
"must_call": ["request_approval", "cancel_order"], "forbid": [], "max_steps": 5},
]
TRIALS = 5
@pytest.mark.parametrize("case", CASES, ids=lambda c: c["id"])
def test_agent_case(case):
passes = 0
for _ in range(TRIALS):
sb = make_sandbox() # isolated state per trial
run = run_agent(case["task"], sb.tools, sb.run_tool, max_steps=case["max_steps"])
called = [t["tool"] for t in run["trajectory"]]
ok = (
all(name in called for name in case["must_call"])
and not any(name in called for name in case["forbid"])
and run["stop_reason"] == "end_turn"
and sb.check_end_state(case["id"]) # assert on the WORLD, not the text
)
passes += ok
assert passes / TRIALS >= 0.8, f"{case['id']}: {passes}/{TRIALS} trials passed"Each trial gets a fresh sandbox so runs cannot contaminate each other. The strongest check is check_end_state: query the fake database to confirm order 2001 is really cancelled, instead of trusting the agent's claim that it is.
How it works
1. Define tasks with verifiable outcomes. Prefer tasks whose success can be checked by code: a value in a database, a file's contents, an API call recorded by a mock. Write the check before you run the agent.
2. Instrument the loop. Your harness records every model turn and tool call (name, arguments, result, error flag), token usage per turn, and the final stop reason. This is the same data observability needs in production, so build it once.
3. Grade at several levels. Outcome checks first; then trajectory rules (required tools, forbidden tools, approval before side effects, no repeated identical calls); then tool-call accuracy (argument validation against the schema plus expected values); then budgets (steps, tokens, time).
4. Repeat and aggregate. Run each task several times; report success rate per task and per slice, and the distribution of steps and tokens (median and worst case matter more than the mean for budgets).
5. Read failed trajectories. Numbers say that something broke; transcripts say why. Categorise failures (wrong tool, bad arguments, gave up early, looped, hallucinated a tool result, ignored a tool error) and fix the most common category first, usually with clearer tool descriptions, better error messages from tools, or prompt guidance.
Trajectory matching strictness is a design choice: exact match (the same calls in the same order) is brittle because many correct paths exist; subsequence / contains (required calls appear in order, extras allowed) is the usual compromise; outcome only is most flexible but misses unsafe routes. Use stricter rules only around safety-critical steps.
task + sandbox ──> agent loop (N trials)
│
records trajectory:
turn ─ tool_use ─ tool_result ─ ...
│
┌──────────────┬───┴──────────┬─────────────┐
v v v v
outcome trajectory tool-call budgets
(end state) (required / accuracy (steps,
forbidden) (name, args) tokens, time)
└──────────────┴──────┬───────┴─────────────┘
v
success rate per task ──> CI gateWhy does it exist?
Agents act in the world: they send messages, change records and spend money with every step. A final-answer check alone will happily pass an agent that deleted the wrong file on the way to a correct summary. Trajectory and budget checks exist to catch unsafe or wasteful behaviour that outcome checks cannot see.
Because agents are stochastic and their behaviour shifts with every prompt or model change, a suite of repeated, sandboxed tasks is the only reliable way to know whether a change made the agent better or worse.
When to use it
For any agent that will run without a human checking each step, and for any change to its system prompt, tool descriptions, tool set or model. Start with 10 to 20 representative tasks including at least a few that require refusing, asking for approval, or recovering from a tool error.
When not to use it
Do not build a heavy agent-eval framework for a fixed workflow (a prompt chain with no model-chosen tools); ordinary input/output tests are enough there. Do not require exact trajectory matches across the board: you will spend your time updating brittle tests instead of improving the agent.
Common mistakes
Grading only the final text, so unsafe or wasteful trajectories pass.
Trusting the agent's claim ('I cancelled the order') instead of checking the end state.
Running each task once and reading a lucky pass as reliability.
Evaluating against real production tools, risking real side effects and flaky results.
Requiring an exact sequence of tool calls when several correct paths exist.
Ignoring step and token budgets until the bill or latency becomes a problem.
Not including error-recovery tasks, so you never learn how the agent handles a failing tool.
Practice exercises
- Easy:
Fix the cancel-order run in the trajectory demo so it passes, then add a third case where the agent calls get_order twice with the same arguments and add a 'no repeated identical calls' check.
- Easy:
List five tasks for an agent you have built, and for each write the end-state check you would use to decide success.
- Medium:
Extend grade() to compute tool-call recall (required calls made / required calls) and argument accuracy separately, and print a per-tool breakdown.
- Medium:
Build a sandbox for a file-editing agent: a temp directory seeded with files, fake tools that operate inside it, and an end-state check that compares files to expected contents.
- Hard:
Run a real agent 10 times on 5 tasks with the Python harness. Report success rate, median and max steps, and median tokens per task. Categorise every failure and fix the most common category.
Interview questions
How is evaluating an agent different from evaluating a single LLM call?
An agent takes many steps and chooses tools, so you must grade the outcome, the path (trajectory) and the resources used (steps, tokens, time), not just one output. It is also more stochastic, so you need repeated trials and success rates rather than single pass/fail results.
What is a trajectory and what would you check in one?
The ordered record of model turns, tool calls with arguments, and tool results. Checks include required tools called (often in order), forbidden tools not called, approval obtained before side effects, valid and correct arguments, no loops or repeated identical calls, and correct handling of tool errors.
Why prefer end-state checks over the agent's final message?
The final message is the agent's claim and can be wrong or hallucinated. Checking the world (database row, file contents, recorded API calls in a mock) verifies what actually happened, and is deterministic.
Explain pass@k versus pass^k.
pass@k is the probability that at least one of k attempts succeeds, appropriate when a human or verifier can pick a good attempt. pass^k is the probability that all k attempts succeed, appropriate for unattended agents that must be consistently reliable. With per-run success r and independence, pass@k = 1 - (1 - r)^k and pass^k = r^k.
How do you keep agent evals safe and repeatable?
Run them in a sandbox: fake or recorded tools, a fresh seeded state per trial, no real credentials or external side effects. Record full trajectories, fix the eval set, and compare versions on the same tasks.
What budgets would you track and why?
Steps per task, input and output tokens (which drive cost), and wall-clock time. Track medians and worst cases. Budgets catch regressions where the agent still succeeds but loops or over-searches, which would hurt cost and user experience in production.