Cost, Latency and Prompt Caching

Tokens drive cost and latency; cut both with prompt caching, batching, streaming, smaller prompts, smaller models and lower effort.

What is it?

LLM APIs bill by tokens: input tokens you send (system prompt, history, retrieved chunks, tool definitions, tool results) and output tokens the model generates. Output tokens are typically priced higher than input tokens and are produced one at a time, so they also dominate generation time. Prices change, so this topic teaches how to reason about cost rather than quoting numbers: always check your provider's current pricing page.

Two latency numbers matter: time to first token (TTFT), how long until the user sees anything, which grows with input size and any thinking the model does before answering; and total time, which adds the time to generate every output token. Users perceive TTFT far more strongly than total time.

The main levers, roughly in order of effort:

  • Send fewer tokens: fewer, better chunks (reranking helps), trimmed chat history (summarise old turns), shorter tool results (return only the fields needed), concise system prompts.
  • Generate fewer tokens: ask for concise answers, set a sensible max_tokens, use structured output instead of prose when code consumes the result.
  • Prompt caching: the provider stores the processed form of a prompt prefix; later requests that start with exactly the same prefix read it from cache, which is cheaper and faster. Cache reads are billed at a fraction of normal input; writing to the cache costs a little more than normal input, so caching pays off when a prefix is reused.
  • Streaming: does not reduce cost or total time, but shows tokens as they are generated, so perceived latency drops to roughly TTFT.
  • Batching: for work nobody is waiting on (nightly evals, bulk classification, contextual retrieval at ingestion), a batch API processes many requests asynchronously at a discount.
  • Right-size the model and effort: route simple steps (classification, routing, extraction, query rewriting) to a smaller, faster model, or lower the effort setting; save the most capable model and higher effort for hard reasoning.
  • Cache answers yourself: an application-level cache keyed on normalised questions (plus the user's permission scope) can skip the LLM entirely for repeated questions.

The prefix rule is the key to prompt caching: the cache matches from the start of the prompt up to a cache breakpoint, in the order tools, then system, then messages. Any change before the breakpoint (a timestamp in the system prompt, a reordered tool list, a different retrieved chunk placed early) means a cache miss. So put large, stable content first and variable content last.

Explain like I'm 10

Prompt caching is like a chef who preps the same base sauce every morning. If every order starts with that sauce, the chef ladles it from the pot instead of cooking it from scratch: faster and cheaper. But if one ingredient at the start of the recipe changes, the whole sauce has to be made again. Streaming is the waiter bringing bread while the main course cooks: the meal takes as long, but the wait feels shorter.

Examples

Where the tokens go: a monthly cost model (runnable)

// Placeholder rates in "units per million tokens" - plug in your provider's current prices.
const RATE = { input: 1, output: 5, cacheRead: 0.1, cacheWrite: 1.25 };
const estimateTokens = text => Math.ceil(text.length / 4); // rough rule of thumb for English

const requestsPerDay = 20000;
function monthlyCost({ system, chunks, chunkTokens, history, question, output, cachedPrefix }) {
  const input = system + chunks * chunkTokens + history + question;
  const cached = cachedPrefix ? system : 0;            // only the stable prefix can be cached
  const perReq = (input - cached) * RATE.input + cached * RATE.cacheRead + output * RATE.output;
  return { inputTokens: input, perMonth: (perReq * requestsPerDay * 30) / 1e6 };
}

const base = { system: 3000, chunks: 10, chunkTokens: 400, history: 1500, question: 30, output: 400, cachedPrefix: false };
const scenarios = {
  "baseline (10 chunks)":            base,
  "rerank -> 4 chunks":              { ...base, chunks: 4 },
  "+ summarise history":             { ...base, chunks: 4, history: 300 },
  "+ cache system prompt":           { ...base, chunks: 4, history: 300, cachedPrefix: true },
  "+ concise answers (250 out)":     { ...base, chunks: 4, history: 300, cachedPrefix: true, output: 250 },
};
const b = monthlyCost(base).perMonth;
for (const [name, s] of Object.entries(scenarios)) {
  const c = monthlyCost(s);
  console.log(name.padEnd(30), "input tok", String(c.inputTokens).padStart(5),
    " cost units", c.perMonth.toFixed(0).padStart(6), " (" + ((c.perMonth / b) * 100).toFixed(0) + "% of baseline)");
}
console.log("chars->tokens example:", estimateTokens("Retrieval-augmented generation grounds answers in your documents."));

With placeholder rates, retrieved chunks are the biggest input cost, so reranking down to fewer, better chunks matters more than shaving the question. Output is a small token count but a high per-token rate, so concise answers also pay off. The chars/4 estimate is only a rough heuristic; measure real counts with the API's usage fields or token-counting endpoint.

The prefix rule: what breaks a prompt cache (runnable)

// A prompt is an ordered list of segments. The cache can reuse the longest identical prefix.
const cache = new Set();
function send(label, segments) {
  let key = "", hitTokens = 0, total = 0, stillHitting = true;
  for (const seg of segments) {
    key += "|" + seg.text;
    total += seg.tokens;
    if (stillHitting && cache.has(key)) hitTokens += seg.tokens; else stillHitting = false;
    cache.add(key);                      // this prefix is now cached for next time
  }
  console.log(label.padEnd(36), "cached", hitTokens, "/", total, "tokens");
}

const TOOLS = { text: "tool definitions", tokens: 1500 };
const DOCS  = { text: "product manual v7", tokens: 20000 };
const q = text => ({ text, tokens: 40 });

// Good: stable content first, the variable question last.
send("good #1 (cold)",               [TOOLS, { text: "system: support bot", tokens: 500 }, DOCS, q("How do I reset?")]);
send("good #2 (new question)",       [TOOLS, { text: "system: support bot", tokens: 500 }, DOCS, q("Is there a warranty?")]);

// Bad: a timestamp at the top of the system prompt changes every request.
send("bad #1 (timestamp in system)", [TOOLS, { text: "system: support bot. Now: 10:00:01", tokens: 510 }, DOCS, q("How do I reset?")]);
send("bad #2 (timestamp in system)", [TOOLS, { text: "system: support bot. Now: 10:00:07", tokens: 510 }, DOCS, q("Is there a warranty?")]);

In the good layout the second request reuses everything except the new question. In the bad layout the tool definitions are still reused, but the changing timestamp breaks the prefix, so the 20,000-token manual after it is processed again every time. Move volatile values (time, user name, retrieved chunks) after the cached content, or into the user message.

Prompt caching and verifying it worked (Python)

import anthropic

client = anthropic.Anthropic()
MANUAL = open("product_manual.txt").read()     # large, stable content

def ask(question: str):
    resp = client.messages.create(
        model="claude-opus-5-5",
        max_tokens=16000,
        system=[
            {"type": "text", "text": "You answer questions about the product manual. Be concise."},
            # Cache breakpoint: everything up to and including this block is the cached prefix.
            {"type": "text", "text": MANUAL, "cache_control": {"type": "ephemeral"}},
        ],
        messages=[{"role": "user", "content": question}],   # variable part goes last
    )
    u = resp.usage
    print(f"input={u.input_tokens} cache_write={u.cache_creation_input_tokens} "
          f"cache_read={u.cache_read_input_tokens} output={u.output_tokens}")
    return resp

ask("How do I reset the device?")     # first call: cache_write > 0, cache_read == 0
ask("What does the warranty cover?")  # within the cache lifetime: cache_read > 0

# Simpler alternative: a top-level cache_control lets the API place the breakpoint
# automatically on the last cacheable block of the request:
# client.messages.create(model="claude-opus-5-5", max_tokens=16000,
#     cache_control={"type": "ephemeral"}, system=..., messages=...)

# Measure a prompt's size before sending it:
count = client.messages.count_tokens(
    model="claude-opus-5-5",
    system=MANUAL,
    messages=[{"role": "user", "content": "How do I reset the device?"}],
)
print("input tokens:", count.input_tokens)

Always verify caching with the usage fields: if cache_read_input_tokens stays 0 on repeat calls, something early in the prompt is changing, or the prefix is below the minimum cacheable length (check the docs for your model). The cache has a short lifetime that is refreshed each time it is read, so it suits steady traffic or bursts of related calls.

Streaming for perceived latency, batches for offline work, low effort for simple steps (Python)

import anthropic

client = anthropic.Anthropic()

# 1) Streaming: the user sees text after time-to-first-token instead of waiting for the whole answer.
with client.messages.stream(
    model="claude-opus-5-5",
    max_tokens=64000,
    messages=[{"role": "user", "content": "Explain our refund policy to a customer."}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)
    final = stream.get_final_message()
print("\noutput tokens:", final.usage.output_tokens)

# 2) Lower effort for a simple, high-volume step such as routing or classification.
resp = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=1000,
    output_config={"effort": "low"},
    messages=[{"role": "user", "content": "Label as billing, technical or other: 'My card was charged twice.'"}],
)

# 3) Message Batches: asynchronous, discounted processing for work nobody is waiting on.
batch = client.messages.batches.create(requests=[
    {"custom_id": f"ticket-{i}",
     "params": {"model": "claude-opus-5-5", "max_tokens": 1000,
                "messages": [{"role": "user", "content": f"Summarise ticket: {t}"}]}}
    for i, t in enumerate(["Login fails on Safari", "Invoice shows wrong VAT"])
])
print(batch.id, batch.processing_status)
# Later (poll until processing_status == "ended"):
# for result in client.messages.batches.results(batch.id):
#     if result.result.type == "succeeded":
#         print(result.custom_id, result.result.message.content[0].text)

Three tools for three situations: streaming for interactive answers, low effort (or a smaller model) for simple high-volume steps, and batches for offline jobs like evals and ingestion. Batch results can arrive in any order, which is why each request carries a custom_id.

How it works

Why input size affects latency. Before generating, the model must process every input token (the prefill). Prefill is parallel and fast per token, but tens of thousands of tokens still add noticeable TTFT. Prompt caching skips recomputing the cached prefix, which is why it reduces latency as well as cost.

Why output dominates total time. Generation is sequential: each new token depends on the previous ones, so generating 1,000 tokens takes roughly ten times as long as 100. Thinking tokens count as output too, so higher effort on hard problems buys quality with time and tokens; lower effort on easy ones saves both.

How caching is billed and verified. On Anthropic's API, the usage block reports cache_creation_input_tokens (written to cache this call), cache_read_input_tokens (served from cache) and input_tokens (the uncached remainder). Total input is the sum of the three. A healthy cached workload shows large cache reads and small uncached input on repeat calls.

Designing prompts for caching: order content from most stable to least stable: tool definitions, system instructions, large reference documents, few-shot examples, conversation history, then the newest user message. In a multi-turn chat, a breakpoint at the end of the conversation lets each turn reuse the previous turns. Keep tool lists in a fixed order and avoid per-request values in the system prompt.

Model and effort routing. A common production shape uses a small, fast model or low effort for routing, query rewriting and classification, and the most capable model for the final grounded answer or complex agent steps. Validate each downgrade on your eval set: if quality holds for that step, keep the savings.

Measure, then optimise. Log tokens (input, cached, output), TTFT and total latency per request and per pipeline step. Optimise the largest bucket first; it is usually retrieved context, long histories or verbose tool results, not the user's question.

request = [tools][system][docs][history][question]
           └──── stable prefix ────┘ └─ varies ─┘
                       │
             cache hit?│yes: read cheap + fast
                       │no : process + write
                       v
  prefill (input) ──> TTFT ──> decode (output, 1 token
                                at a time) ──> done
  stream: user sees text from TTFT onward

Why does it exist?

At prototype scale cost and latency are invisible; at production scale they decide whether a feature is viable. A RAG answer that sends ten unnecessary chunks to the largest model at the highest effort can cost many times more and take seconds longer than it needs to, with no quality gain. These techniques exist to spend tokens only where they improve results.

When to use it

Cache whenever a large prefix repeats across requests (system prompt plus reference docs, long agent tool lists, multi-turn chats). Stream every user-facing answer. Batch every offline job. Route simple steps to cheaper settings once your evals confirm quality holds. Revisit after any prompt change, since a single moved line can break caching.

When not to use it

Do not cache prompts that are unique per request or used rarely; you pay the write premium without reads. Do not use batches for anything a user is waiting on. Do not cut chunks, history or effort below the point where your eval metrics drop: a cheap wrong answer is the most expensive kind. And do not optimise from guesses; measure the token breakdown first.

Common mistakes

  • Putting timestamps, request ids or user names at the top of the system prompt, breaking the cache prefix on every call.

  • Assuming caching works without checking cache_read_input_tokens in the usage data.

  • Reordering or regenerating tool definitions per request, which changes the earliest part of the prefix.

  • Streaming in the UI but buffering the whole response in a proxy or server in between.

  • Sending entire documents or full tool payloads when a few fields would do.

  • Using the largest model and highest effort for trivial classification and routing steps.

  • Estimating tokens from word counts for billing decisions instead of using real usage numbers.

Practice exercises

  1. Easy:

    In the cost model demo, find which single change saves the most and explain why. Then add a scenario for 'answers capped at 150 tokens'.

  2. Easy:

    Rewrite this layout for caching: [system with today's date][tools][user question][product docs].

  3. Medium:

    Instrument a RAG pipeline to log input, cache-read, cache-write and output tokens per request, plus TTFT and total time, and produce a daily summary.

  4. Medium:

    Move your eval suite's judge calls to the Message Batches API and compare wall-clock time and cost against synchronous calls.

  5. Hard:

    Build a two-tier pipeline: a low-effort routing/rewriting step and a high-capability answer step. Show on your eval set that quality holds while tokens and latency drop, or find the step where it does not.

Interview questions

What determines the cost of an LLM call?

Token counts multiplied by per-token prices: input tokens (system prompt, history, retrieved context, tool definitions and results) and output tokens (including thinking), with output usually priced higher. Cached input is billed differently: cache reads cost less than normal input, cache writes slightly more.

Explain how prompt caching works and its main pitfall.

The provider stores the processed state of a prompt prefix up to a breakpoint; later requests with an identical prefix reuse it, cutting cost and time to first token. It is an exact prefix match in the order tools, system, messages, so any change earlier in the prompt (a timestamp, reordered tools, a different chunk) invalidates everything after it. Verify with cache read token counts.

Does streaming reduce cost or latency?

It reduces perceived latency: the user starts reading after the time to first token instead of waiting for the full response. It does not reduce token cost or total generation time.

When would you use a batch API?

For workloads with no user waiting: evaluations, bulk classification or summarisation, ingestion-time enrichment like contextual retrieval. Batches are asynchronous and discounted, and results are matched back with custom ids.

How do you decide whether a smaller model or lower effort is acceptable for a step?

Measure it: run the step's eval set (routing accuracy, extraction correctness, end-to-end answer quality) with both settings. If quality holds within your tolerance, take the savings; simple, well-specified steps usually qualify, open-ended reasoning often does not.

What is time to first token and what affects it?

The delay before the first output token arrives. It grows with input size (prefill), with any thinking done before visible output, and with queueing. Prompt caching, smaller inputs and lower effort reduce it; streaming makes it the latency users actually perceive.