Observability and Debugging LLM Apps

Trace every model and tool call with inputs, outputs, tokens and latency; log structurally; classify failures and debug from traces.

What is it?

Observability means being able to answer 'what happened, and why?' about any request after the fact, from the data your system recorded. For LLM apps this matters even more than for ordinary software: the same input can produce different outputs, failures are often wrong answers rather than exceptions, and a single user request may fan out into retrieval, several model calls and many tool calls.

The core building block is the trace: one record per user request, made of nested spans, one per operation (retrieve, rerank, LLM call, tool call). Each span records:

  • Inputs: the full prompt (system, messages, tool definitions version), retrieved chunk ids, tool arguments.
  • Outputs: the response content, tool results, stop reason.
  • Usage: input, cached and output tokens, from which you compute cost.
  • Timing: start, end, duration; for streamed calls also time to first token.
  • Metadata: model id, effort setting, prompt version, user or tenant id (pseudonymised), request id, errors and retries.

Logs should be structured (one JSON object per event with consistent field names), not free text, so you can query them: 'all requests yesterday where stop_reason was max_tokens', 'p95 latency of the rerank span', 'cost per tenant'. Link every span by a shared trace_id.

Debugging LLM apps is mostly reading traces. When a user reports a bad answer you open its trace and walk it: was the question rewritten sensibly? Did retrieval return the right chunks? Was the right context actually in the final prompt? Did the model ignore it, or was it truncated by max_tokens? Did a tool fail and get retried? This pinpoints the broken stage.

Over time you build a failure taxonomy, a fixed set of categories such as retrieval miss, ignored context, hallucinated detail, wrong tool, bad tool arguments, tool error not recovered, truncated output, refusal, format error, timeout. Tagging failures with these categories turns anecdotes into counts and tells you what to fix first. Every analysed failure should also become a new case in your golden dataset.

Standards and tools exist for this: OpenTelemetry is the common open standard for traces and has emerging conventions for generative-AI spans, and many LLM-specific observability products and open-source tools build on similar ideas. The concepts in this topic apply whichever you choose.

Explain like I'm 10

A trace is like a flight recorder for each request. When something goes wrong you do not guess; you replay the recording: what the pilot saw (inputs), what they did (model and tool calls), how long each step took, and where it went off course. The failure taxonomy is the accident-report form that makes reports comparable, so patterns become visible.

Examples

A minimal tracer around a RAG request (runnable)

// Fake clock so output is deterministic; in real code use performance.now() or Date.now().
let clock = 0;
const now = () => clock;
const spend = ms => { clock += ms; };

const spans = [];
function span(traceId, name, parent, fn) {
  const s = { traceId, id: "s" + (spans.length + 1), parent, name, start: now() };
  spans.push(s);
  try {
    const result = fn(s);
    s.status = "ok";
    return result;
  } catch (e) {
    s.status = "error"; s.error = e.message; throw e;
  } finally {
    s.ms = now() - s.start;
  }
}

// Fake pipeline pieces.
function retrieve(q) { spend(40); return ["refunds.md#0", "billing.md#3"]; }
function llm(prompt) { spend(900); return { text: "Refunds are allowed within 30 days [refunds.md#0].",
  usage: { input: 1850, cacheRead: 1200, output: 22 }, stop: "end_turn" }; }

function handle(question) {
  const traceId = "t-001";
  return span(traceId, "request", null, root => {
    root.input = question;
    const ids = span(traceId, "retrieve", root.id, s => { const r = retrieve(question); s.output = r; return r; });
    const resp = span(traceId, "llm.answer", root.id, s => {
      s.model = "claude-opus-5-5"; s.promptVersion = "answer-v3"; s.chunkIds = ids;
      const r = llm("...prompt with " + ids.join(",") + "...");
      s.usage = r.usage; s.stopReason = r.stop; s.output = r.text;
      return r;
    });
    root.output = resp.text;
    return resp.text;
  });
}

handle("What is the refund window?");
for (const s of spans) {
  const indent = s.parent ? "  " : "";
  console.log(indent + JSON.stringify({ trace: s.traceId, span: s.name, ms: s.ms, status: s.status,
    usage: s.usage, stop: s.stopReason, out: s.chunkIds || s.output }));
}
const totalTokens = spans.filter(s => s.usage).reduce((a, s) => a + s.usage.input + s.usage.output, 0);
console.log("total tokens:", totalTokens, " slowest span:", spans.slice(1).sort((a, b) => b.ms - a.ms)[0].name);

Every operation becomes a span with timing, inputs, outputs and usage, linked by trace id and parent id. Note the request span ends last but is printed first because it started first. From these records you can answer: which chunks did the model see, which prompt version produced this answer, where did the time go?

From logged failures to a prioritised failure taxonomy (runnable)

// Each reviewed failure is tagged with exactly one category (plus notes).
const reviewed = [
  { trace: "t-101", category: "retrieval_miss", note: "SSO docs not indexed" },
  { trace: "t-102", category: "ignored_context", note: "answered from general knowledge" },
  { trace: "t-103", category: "retrieval_miss", note: "query used old product name" },
  { trace: "t-104", category: "truncated_output", note: "stop_reason max_tokens" },
  { trace: "t-105", category: "bad_tool_args", note: "date passed as 'next friday'" },
  { trace: "t-106", category: "retrieval_miss", note: "answer split across two pages" },
  { trace: "t-107", category: "hallucinated_detail", note: "invented a 24h SLA" },
  { trace: "t-108", category: "bad_tool_args", note: "order id missing leading zeros" },
];
const FIX_HINT = {
  retrieval_miss: "chunking / hybrid search / query rewriting / index coverage",
  ignored_context: "prompt: answer only from documents; put docs before question",
  truncated_output: "raise max_tokens or ask for shorter answers",
  bad_tool_args: "tighter input_schema, better descriptions, validate + return clear is_error",
  hallucinated_detail: "grounding instructions, citations, faithfulness eval",
};
const counts = {};
for (const r of reviewed) counts[r.category] = (counts[r.category] || 0) + 1;
Object.entries(counts).sort((a, b) => b[1] - a[1]).forEach(([cat, n]) => {
  const pct = ((n / reviewed.length) * 100).toFixed(0);
  console.log((cat + " ").padEnd(20, "."), String(n).padStart(2), "(" + pct + "%)  fix:", FIX_HINT[cat]);
});

Counting categories turns a pile of complaints into a priority list: here retrieval misses dominate, so improving retrieval beats any amount of prompt polishing. Keep the category list small and fixed so counts stay comparable week to week.

Tracing wrapper for real API calls with structured JSON logs (Python)

import json
import logging
import time
import uuid
import anthropic

client = anthropic.Anthropic()
log = logging.getLogger("llm")
logging.basicConfig(level=logging.INFO, format="%(message)s")

def redact(text: str) -> str:
    # Replace with real PII scrubbing (emails, phone numbers, account ids) before logging.
    return text

def traced_create(trace_id: str, span_name: str, prompt_version: str, **kwargs):
    span_id = uuid.uuid4().hex[:8]
    start = time.perf_counter()
    record = {"trace_id": trace_id, "span_id": span_id, "span": span_name,
              "model": kwargs.get("model"), "prompt_version": prompt_version,
              "effort": (kwargs.get("output_config") or {}).get("effort"),
              "n_messages": len(kwargs.get("messages", [])),
              "tools": [t["name"] for t in kwargs.get("tools", [])]}
    try:
        resp = client.messages.create(**kwargs)
    except anthropic.APIError as e:
        record.update(status="error", error_type=type(e).__name__, error=str(e)[:500],
                      ms=round((time.perf_counter() - start) * 1000))
        log.error(json.dumps(record))
        raise
    u = resp.usage
    record.update(
        status="ok",
        ms=round((time.perf_counter() - start) * 1000),
        message_id=resp.id,
        stop_reason=resp.stop_reason,
        input_tokens=u.input_tokens,
        cache_read_tokens=u.cache_read_input_tokens,
        cache_write_tokens=u.cache_creation_input_tokens,
        output_tokens=u.output_tokens,
        tool_calls=[{"name": b.name, "input": b.input} for b in resp.content if b.type == "tool_use"],
        output_preview=redact("".join(b.text for b in resp.content if b.type == "text"))[:500],
    )
    log.info(json.dumps(record))
    if resp.stop_reason == "max_tokens":
        log.warning(json.dumps({"trace_id": trace_id, "span_id": span_id, "alert": "truncated_output"}))
    return resp

# Usage: one trace id per user request, shared by every span in it.
trace_id = uuid.uuid4().hex
resp = traced_create(trace_id, "answer", "answer-v3",
    model="claude-opus-5-5", max_tokens=16000,
    messages=[{"role": "user", "content": "What is the refund window?"}])

Every call emits one JSON line with the fields you will want to query later: tokens (including cache), latency, stop reason, tool calls, prompt version and the shared trace id. Store full prompts and outputs in a separate, access-controlled store if your privacy policy allows it, and redact personal data before it reaches general logs.

How it works

Instrument at the boundaries. Wrap the few places where work happens: the LLM client call, each tool execution, retrieval and reranking, and the request handler. Pass a trace id through the request (as a function argument or context variable) so every span links back. If you use OpenTelemetry, each of these becomes a span with attributes; the same fields apply.

Version everything that shapes output: prompt templates (a name and version in the span), tool definitions, retrieval settings (k, reranker on/off), model id and effort. Without versions you cannot correlate a quality drop with the change that caused it.

Derive metrics from traces: requests per minute, error rate by type (rate limits, overloaded, timeouts), p50/p95 latency per span, TTFT for streamed responses, tokens and cost per request, per feature and per tenant, cache hit ratio (cache read tokens / total input tokens), stop-reason distribution (a rise in max_tokens signals truncation), tool error rate, steps per agent run. Alert on sudden changes.

Quality signals in production: explicit user feedback (thumbs up/down with trace id attached), implicit signals (user rephrases immediately, copies the answer, escalates to a human), and sampled online evaluation (run the LLM judge on a small percentage of traffic). Low-rated traces go into a review queue.

The debugging loop: pick failing traces (from feedback, alerts or sampled evals) → read the trace stage by stage → tag the failure category → reproduce it locally by replaying the recorded inputs → add it to the golden dataset → fix → confirm with the eval suite that it is fixed and nothing else regressed.

Privacy and retention: prompts and outputs often contain personal or confidential data. Decide what to store, redact before logging where possible, restrict access to raw traces, and set retention periods. Observability must not become a data leak.

trace t-001  (user request)
├─ span rewrite_query   120ms  tok 300/40
├─ span retrieve         45ms  chunks [a,b,c]
├─ span rerank           80ms  kept [a,c]
├─ span llm.answer      950ms  tok 2100/180
│     stop=tool_use  prompt=answer-v3
├─ span tool.read_file   15ms  ok
└─ span llm.answer      700ms  stop=end_turn
        │
        v
 JSON logs ──> metrics, alerts, review queue
        └──> failures tagged ──> golden set

Why does it exist?

LLM failures are usually silent: the HTTP call succeeds and returns a fluent but wrong answer. Without traces you cannot tell whether retrieval, the prompt, the model or a tool caused it, and without token tracking you discover cost problems on the invoice. Observability exists to make these invisible failures visible and diagnosable.

When to use it

From the first deployment, even internal ones. Tracing is cheap to add early and very hard to retrofit once you have a production incident and no data. At minimum, log per call: trace id, model, prompt version, tokens, latency, stop reason, tool calls and errors.

When not to use it

Do not log full prompts and outputs containing personal or regulated data into general-purpose logs without redaction, access control and retention rules. Do not build a bespoke observability platform if an existing tracing standard or tool fits; the value is in using the data, not in building the plumbing.

Common mistakes

  • Logging only errors, when most LLM failures are successful calls with wrong content.

  • Free-text log lines that cannot be queried or aggregated.

  • No trace id linking retrieval, model and tool spans of the same request.

  • Not recording prompt versions, so quality changes cannot be tied to the change that caused them.

  • Ignoring stop_reason, missing silent truncation at max_tokens.

  • Tracking total spend but not tokens per feature or tenant, so cost spikes cannot be attributed.

  • Logging raw personal data with no redaction or retention policy.

  • Fixing reported failures without adding them to the eval set, so they quietly return later.

Practice exercises

  1. Easy:

    Extend the runnable tracer with a 'rerank' span between retrieve and llm.answer, and make the llm call throw once to see an error span.

  2. Easy:

    Write down the fields you would log for every tool call in an agent and justify each.

  3. Medium:

    Wrap your agent loop with traced_create and a traced tool runner sharing one trace id. Produce a per-request summary: steps, tokens, cost estimate, slowest span.

  4. Medium:

    Write a script that reads a day of JSON logs and reports p50/p95 latency per span, stop-reason distribution and cache hit ratio.

  5. Hard:

    Review 30 real or simulated failed traces, tag them with a taxonomy of at most 8 categories, fix the top category, add the 30 cases to your golden set and show the before/after eval results.

Interview questions

What would you record for every LLM call in production?

Trace and span ids, model id, effort or other settings, prompt template version, inputs (or a reference to them in a protected store), output and tool calls, stop reason, input/cache-read/cache-write/output tokens, latency and TTFT, retries and errors, and a pseudonymised user or tenant id.

How do you debug a user's report of a wrong answer?

Find the request's trace, then walk the stages: was the query rewritten sensibly, did retrieval return relevant chunks, were they in the final prompt, did the model use or ignore them, was the output truncated, did any tool fail? Tag the failure category, reproduce by replaying the inputs, add the case to the golden set, fix, and confirm with evals.

What is a failure taxonomy and why keep one?

A small fixed set of failure categories (retrieval miss, ignored context, hallucination, wrong tool, bad arguments, truncation, refusal, timeout, and so on). Tagging failures consistently turns anecdotes into counts so you fix the most frequent problems first and can see whether fixes work over time.

Which production metrics are specific to LLM apps?

Tokens and cost per request and per feature, cache hit ratio, TTFT, stop-reason distribution (especially max_tokens), tool error rate, agent steps per task, refusal rate, user feedback rates, and sampled online eval scores such as faithfulness.

How do you balance observability and privacy?

Redact or pseudonymise personal data before logging, keep raw prompts and outputs in an access-controlled store with a retention limit, log only what is needed, and document it. Aggregated metrics and structural fields (tokens, latency, chunk ids) rarely need raw content.