Tokens and Context Windows
What tokens are, how tokenizers split text, why everything is counted in tokens, and how the context window limits what a model can see.
What is it?
Models do not read letters or words; they read tokens. A token is a chunk of text from a fixed vocabulary that the model's tokenizer knows. Common words are often a single token (the, cat), rarer words are split into pieces (tokenization might become token + ization), and spaces and punctuation are part of tokens too. Each token maps to an integer token id.
Why not just use characters or whole words? Characters make sequences very long (more work per sentence). Whole words make the vocabulary enormous and cannot handle new words, typos or code. Subword tokenizers are the compromise: frequent chunks get their own token, and anything else can still be spelled out from smaller pieces. The most common family is Byte Pair Encoding (BPE): start from single bytes/characters and repeatedly merge the most frequent adjacent pair into a new token.
A rough rule of thumb for English prose is that a token is about three-quarters of a word, or around four characters. Treat this only as a quick estimate: code, numbers, non-English languages and unusual formatting usually use more tokens per character, and different model families use different tokenizers. When the number matters, count with the provider's tokenizer or token-counting endpoint.
Tokens matter for three practical reasons:
- Cost: LLM APIs bill by input tokens (what you send) and output tokens (what the model writes), usually at different rates.
- Speed: output is produced one token at a time, so longer answers take proportionally longer.
- Limits: every model has a maximum context window and you set a maximum output length.
The context window is the maximum number of tokens the model can consider in one request - your system prompt, the whole conversation history, any documents you paste in, tool definitions, and the tokens it generates in its reply all share it. Anything outside the window simply does not exist for the model. Context windows of modern models are large (whole books can fit), but they are finite, and filling them has costs: more money, more latency, and often worse attention to details buried in the middle of a huge prompt.
Because the API is stateless (it does not remember previous requests), a chat application resends the entire conversation each turn. Long conversations therefore grow until they approach the window and must be trimmed or summarised. This is the root of many design decisions in RAG and agents: you cannot paste everything, so you must choose what goes into the context.
Tokenization also explains some odd model behaviours: counting the letters in a word, reversing strings, or exact character-level arithmetic can be harder than expected, because the model sees straw + berry as chunks, not as individual letters.
Explain like I'm 10
Tokens are like Scrabble tiles that can hold several letters at once: common chunks get their own tile ('ing', 'the'), rare words must be built from several smaller tiles. The context window is the size of the table: the model can only look at the tiles on the table right now. Your instructions, the chat so far, any documents and the reply being written all have to fit on that table together.
Examples
A toy tokenizer: greedy longest-match against a vocabulary
// Real tokenizers have tens of thousands of entries learned from data.
const vocab = ["un", "believ", "able", "token", "ization", "the", " ", "s",
"a", "b", "c", "d", "e", "i", "l", "n", "o", "r", "t", "u", "v", "z", "k", "!"];
const idOf = Object.fromEntries(vocab.map((t, i) => [t, i]));
function tokenize(text) {
const tokens = [];
let i = 0;
while (i < text.length) {
// try the longest vocabulary entry that matches at position i
let match = null;
for (const t of vocab) {
if (text.startsWith(t, i) && (!match || t.length > match.length)) match = t;
}
if (!match) throw new Error("no token for '" + text[i] + "'");
tokens.push(match);
i += match.length;
}
return tokens;
}
for (const text of ["unbelievable", "tokenization", "the tokens", "unbelievable!"]) {
const toks = tokenize(text);
console.log(JSON.stringify(text), "->", toks.map(t => JSON.stringify(t)).join(" "),
"| ids:", toks.map(t => idOf[t]).join(","), "| count:", toks.length);
}'unbelievable' costs 3 tokens because the vocabulary has those chunks. 'the tokens' needs a space token and then falls back to 's'. Words the vocabulary covers well are cheap; unusual strings cost more tokens.
How a BPE vocabulary is learned: merge the most frequent pair
// Start with characters; repeatedly merge the most frequent adjacent pair.
const words = ["low", "low", "low", "lower", "lowest", "newest", "newest", "widest"];
let seqs = words.map(w => w.split(""));
function countPairs(seqs) {
const counts = new Map();
for (const s of seqs) {
for (let i = 0; i < s.length - 1; i++) {
const pair = s[i] + "|" + s[i + 1];
counts.set(pair, (counts.get(pair) || 0) + 1);
}
}
return counts;
}
for (let step = 1; step <= 6; step++) {
const counts = countPairs(seqs);
let best = null, bestCount = 0;
for (const [pair, c] of counts) if (c > bestCount) { best = pair; bestCount = c; }
const [a, b] = best.split("|");
seqs = seqs.map(s => {
const out = [];
for (let i = 0; i < s.length; i++) {
if (i < s.length - 1 && s[i] === a && s[i + 1] === b) { out.push(a + b); i++; }
else out.push(s[i]);
}
return out;
});
console.log("merge " + step + ": '" + a + "' + '" + b + "' -> '" + a + b + "' (seen " + bestCount + "x)");
}
const unique = [...new Set(words)];
for (const w of unique) {
console.log(w.padEnd(7), "->", seqs[words.indexOf(w)].join(" | "));
}After a few merges, frequent chunks like 'low' and 'est' become single tokens. Real tokenizers run tens of thousands of merges over huge corpora, usually on raw bytes so any text (emoji, any language, binary-ish strings) can be encoded.
Budgeting a context window
// Rough estimate only: ~4 characters per token for English prose.
const estimateTokens = text => Math.ceil(text.length / 4);
const WINDOW = 200000; // example number for this exercise, not a real model's limit
const MAX_OUTPUT = 16000; // reserve room for the reply
const systemPrompt = "You are a helpful support assistant for Acme Cloud.";
const history = Array.from({ length: 40 }, (_, i) => "turn " + i + ": " + "x".repeat(2000));
const document = "y".repeat(700000);
const parts = {
system: estimateTokens(systemPrompt),
history: history.reduce((s, h) => s + estimateTokens(h), 0),
document: estimateTokens(document),
};
const input = parts.system + parts.history + parts.document;
console.log("estimated input tokens:", input, parts);
console.log("input + reserved output:", input + MAX_OUTPUT, "of", WINDOW);
if (input + MAX_OUTPUT > WINDOW) {
// Strategy: keep the system prompt, keep the newest turns, drop the oldest
let budget = WINDOW - MAX_OUTPUT - parts.system - parts.document;
const kept = [];
for (let i = history.length - 1; i >= 0; i--) {
const t = estimateTokens(history[i]);
if (t > budget) break;
kept.unshift(history[i]);
budget -= t;
}
console.log("too big: keeping the newest", kept.length, "of", history.length, "turns");
}Every application that keeps a conversation going needs a policy like this: reserve output space, keep instructions, and trim or summarise old history. Better still, do not paste whole documents - retrieve only the relevant parts (what-is-rag).
Counting tokens exactly with the API (Python)
import anthropic
client = anthropic.Anthropic()
messages = [{"role": "user", "content": "Summarise the plot of Hamlet in three sentences."}]
# Count before sending (useful for budgeting and trimming)
count = client.messages.count_tokens(
model="claude-opus-5-5",
system="You are a concise literature tutor.",
messages=messages,
)
print("input tokens (counted):", count.input_tokens)
resp = client.messages.create(
model="claude-opus-5-5",
max_tokens=16000,
system="You are a concise literature tutor.",
messages=messages,
)
print("input tokens (billed):", resp.usage.input_tokens)
print("output tokens:", resp.usage.output_tokens)
print("stop reason:", resp.stop_reason) # "max_tokens" means the reply was cut offNever guess token counts in production code that enforces limits; use the provider's counter. Every response also reports actual usage, which you should log for cost tracking.
How it works
Tokenizer training happens once, before the model is trained: an algorithm such as BPE scans a large corpus and builds a vocabulary of frequent chunks. That vocabulary is then fixed for the model's lifetime - which is why different model families count tokens differently.
At request time, your text is converted to token ids, the model processes them, and the generated ids are converted back to text (detokenised). The API reports how many input and output tokens were used.
The context window is a hard limit of the model architecture and serving setup: the total of input tokens plus generated output tokens must fit. max_tokens in your request caps the output part. If the model reaches it, generation stops mid-thought and the response's stop reason says so (max_tokens).
Cost of long context: attention compares tokens with each other, so processing very long inputs takes more compute and time. Providers price by token, and some offer prompt caching so a large, unchanging prefix (like a long document or system prompt) is cheaper and faster on repeat requests (see cost-and-latency).
Quality in long context: models can usually find information anywhere in their window, but performance on detailed reasoning tends to be best when the context is focused and relevant. Placing long documents first and the question at the end, and clearly labelling sections, helps.
CONTEXT WINDOW (fixed size)
+-------------------------------------------------+
| system | tools | history ...| docs | question | |
|<------------- input tokens ------------->|<out>|
+-------------------------------------------------+
^
reply grows here, token by
token, up to max_tokens
"tokenization!" -> [token][ization][!] -> [8, 912, 0]Why does it exist?
Neural networks work on numbers, so text must be turned into numbers. Tokens are the unit that balances vocabulary size against sequence length. The context window exists because memory and compute grow with sequence length; providers set a maximum that their hardware and model can handle well.
When to use it
Think in tokens whenever you estimate cost, set max_tokens, decide how much history to keep, choose chunk sizes for RAG, or debug a cut-off answer. Count tokens precisely (with the API) when you enforce hard limits; use the rough 4-characters estimate only for back-of-the-envelope planning.
When not to use it
Do not try to fill the context window just because it is large. Stuffing in every document is slower, more expensive and often less accurate than retrieving the few relevant passages. Do not hard-code token estimates from one model's tokenizer when switching models; re-measure.
Common mistakes
Assuming one word equals one token; code, numbers and non-English text often use many more.
Forgetting that the reply's tokens also count toward the context window, then getting truncated answers.
Setting
max_tokenstoo low and not checkingstop_reason == "max_tokens".Letting chat history grow forever until requests fail or become slow and expensive.
Pasting whole knowledge bases into the prompt instead of retrieving relevant parts.
Using a character-based estimate for billing or hard limits instead of the provider's token counter.
Expecting reliable character-level operations (counting letters, exact reversals) without giving the model a tool or code to do it.
Practice exercises
- Easy:
Using the toy tokenizer, add two vocabulary entries that reduce the token count of 'the tokens' and show the before/after counts.
- Easy:
Explain why the context window must hold both the conversation history and the model's reply.
- Medium:
Run the BPE demo with 12 merges instead of 6. Which tokens appear, and how many tokens does each word use at the end?
- Medium:
Write a function
trimHistory(messages, budget)that keeps the first (system-like) message and as many of the most recent messages as fit in the token budget, using the rough estimator. - Hard:
Build a mini BPE tokenizer with two functions:
train(corpus, numMerges)returns an ordered list of merges, andencode(text, merges)applies them in order to a new string. Test that encoding a training word yields the same tokens as training produced.
Interview questions
What is a token?
A unit of text from the model's fixed vocabulary - a word, sub-word, punctuation mark or byte sequence - represented as an integer id. Models read and write tokens, and APIs bill by them.
Why do LLMs use sub-word tokenization?
It balances vocabulary size and sequence length: frequent chunks are single tokens, while rare words, typos, code and any language can still be represented by combining smaller pieces, so nothing is out of vocabulary.
What is a context window?
The maximum number of tokens a model can handle in one request, covering system prompt, history, documents, tool definitions and the generated output together.
A reply is cut off mid-sentence. What do you check?
The stop_reason: if it is max_tokens, the output limit was hit. Raise max_tokens, ask for a shorter answer, or continue in a follow-up request. Also check the total context is not near the window limit.
How do you handle a conversation that is getting too long?
Keep the system prompt, keep recent turns, and drop or summarise older turns; store important facts separately and retrieve them when needed. Count tokens with the provider's counter to enforce the budget.
Why might an LLM fail to count the letters in a word?
It sees tokens, not characters. A word may be one or two tokens, so the individual letters are not directly visible; character-level tasks are better done with code or a tool.