How Large Language Models Work

Next-token prediction, training data, the transformer and attention, pretraining vs instruction tuning vs RLHF, and why models can be wrong.

What is it?

A Large Language Model does one thing: given some text, it predicts what comes next. Everything else - answering questions, writing code, following instructions - is built on top of that single ability.

Text is first split into tokens, small chunks such as whole words, parts of words or punctuation (un, believ, able). Each token has an id number. The model sees a sequence of token ids and outputs a probability distribution over its whole vocabulary for the next token. A probability distribution is just a list of possible outcomes with a likelihood for each, where the likelihoods add up to 1 (100%). For the input 'The cat sat on the', the model might output: mat 0.41, floor 0.18, sofa 0.09, ... thousands more with tiny probabilities.

To generate text, the system picks one token from that distribution, appends it to the input, and asks the model again. Token by token, one at a time. This is called autoregressive generation. When you watch a chatbot 'type', you are literally watching this loop.

Where does the knowledge come from? In pretraining, the model reads a huge amount of text (web pages, books, code, articles). For every position in that text it tries to predict the next token, compares its guess with the real next token, and adjusts its parameters slightly using gradient descent (see what-is-ai). After trillions of such guesses, the parameters encode grammar, facts, reasoning patterns and styles - because predicting text well requires those things. To predict the end of 'The capital of France is', it helps to have learned that the answer is Paris.

The transformer. Modern LLMs use a neural network architecture called the transformer. Its key idea is attention: when processing each token, the model looks back at all the earlier tokens and decides how much each one matters for understanding this one. In 'The animal didn't cross the street because it was too tired', attention lets the model connect it to animal. Each token is represented as a list of numbers (a vector); attention computes how well each pair of tokens 'matches' and blends information from the best matches. A transformer stacks dozens of these attention layers, each refining the representation.

From text predictor to assistant. A model that has only been pretrained (a base model) continues text; ask it 'What is the capital of France?' and it might continue with more quiz questions, because that is what such text looks like on the web. Two more training stages turn it into a helpful assistant:

  • Instruction tuning (supervised fine-tuning): further training on many examples of instructions paired with good responses, written or curated by people. The model learns the format of being an assistant: read the request, answer it.
  • Reinforcement learning from human feedback (RLHF) and related methods: the model produces several answers, people (or a model trained on people's preferences) rank which is better, and the model is trained to produce more of the preferred kind - more helpful, more honest, safer. Some labs also use AI feedback guided by a written set of principles.

Why models can be wrong. An LLM is optimised to produce plausible continuations, not verified true ones. It has no built-in database of facts and no automatic way to check its output. If the training data was thin, outdated or contradictory on a topic, the most plausible-sounding continuation may be false - this is a hallucination. Its knowledge also stops at a training cutoff date, so it does not know about later events unless you give it that information in the prompt.

What the model does NOT do: it does not search the internet by itself (unless given a tool), it does not remember your previous conversations (unless you send them again), and it does not learn from your chats in real time - its parameters are fixed during inference.

Explain like I'm 10

Imagine the world's most voracious reader playing a game: you read them the start of a sentence and they must guess the next word. After reading a library bigger than any human could, they get astonishingly good - good enough that 'guessing the next word' of a question produces a correct, well-written answer. Instruction tuning is a finishing school that teaches them to respond to requests instead of rambling on. But they are still guessing: if they never read about something, they may confidently guess wrong rather than say 'I don't know'.

Examples

A toy next-word predictor from bigram counts

// A "bigram model": predict the next word from the current word only,
// using counts from a tiny training corpus. LLMs use the WHOLE preceding
// context and billions of parameters, but the job is the same.
const corpus = [
  "the cat sat on the mat",
  "the cat ate the fish",
  "the dog sat on the rug",
  "the dog chased the cat",
];

// Training: count which word follows which
const next = {};
for (const sentence of corpus) {
  const words = sentence.split(" ");
  for (let i = 0; i < words.length - 1; i++) {
    const a = words[i], b = words[i + 1];
    next[a] = next[a] || {};
    next[a][b] = (next[a][b] || 0) + 1;
  }
}

// Inference: turn counts into a probability distribution
function predict(word) {
  const options = next[word] || {};
  const total = Object.values(options).reduce((s, c) => s + c, 0);
  return Object.entries(options)
    .map(([w, c]) => [w, c / total])
    .sort((x, y) => y[1] - x[1]);
}

for (const w of ["the", "cat", "sat"]) {
  const dist = predict(w).map(([t, p]) => t + "=" + p.toFixed(2)).join(", ");
  console.log("after '" + w + "':", dist);
}

// Autoregressive generation: always take the most likely word (greedy)
let word = "the";
const out = [word];
for (let i = 0; i < 5; i++) {
  const dist = predict(word);
  if (dist.length === 0) break;
  word = dist[0][0];
  out.push(word);
}
console.log("generated:", out.join(" "));

Notice the generated text loops ('the cat sat on the cat ...') because a bigram model only sees one word of context. Transformers fix this with attention over the entire context, which is why they produce coherent paragraphs.

Attention intuition: which earlier word does 'it' look at?

// Each token gets a small vector (real models: thousands of numbers).
// Attention = compare the current token's QUERY with every token's KEY,
// turn the scores into weights with softmax, and blend.
const tokens = ["the", "animal", "crossed", "the", "street", "because", "it"];
const keys = {
  the:     [0.0, 0.1, 0.0],
  animal:  [0.9, 0.8, 0.1],
  crossed: [0.1, 0.0, 0.9],
  street:  [0.2, 0.9, 0.0],
  because: [0.0, 0.0, 0.3],
  it:      [0.6, 0.4, 0.1],
};
const query = [3.0, 1.5, 0.0]; // what "it" is looking for: a living thing / a noun

const dot = (a, b) => a.reduce((s, x, i) => s + x * b[i], 0);
const scale = Math.sqrt(query.length);
const scores = tokens.map(t => dot(query, keys[t]) / scale);

// softmax: exponentiate and normalise so weights sum to 1
const exps = scores.map(s => Math.exp(s));
const sum = exps.reduce((a, b) => a + b, 0);
const weights = exps.map(e => e / sum);

tokens.forEach((t, i) => {
  const bar = "#".repeat(Math.round(weights[i] * 60));
  console.log(t.padEnd(8), weights[i].toFixed(3), bar);
});

The vectors here are hand-picked to show the idea. In a real transformer, the query and key vectors are computed from the tokens by learned parameters, and many attention 'heads' run in parallel, each learning to track a different kind of relationship (who did what, which noun a pronoun refers to, matching brackets in code...).

Watching autoregressive generation happen (Python, streaming)

import anthropic

client = anthropic.Anthropic()

# Streaming shows text as it is generated, chunk by chunk.
with client.messages.stream(
    model="claude-opus-5-5",
    max_tokens=64000,
    messages=[{"role": "user", "content": "Write one sentence about why the sky is blue."}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)

final = stream.get_final_message()
print()
print("output tokens generated:", final.usage.output_tokens)

Each printed chunk corresponds to one or a few newly generated tokens. The output token count tells you how many times the model ran its next-token step.

How it works

Step 1 - Tokenise. The input text becomes a list of token ids, e.g. [464, 3797, 3332, 319, 262].

Step 2 - Embed. Each token id is looked up in a learned table that maps it to a vector (a list of numbers). Information about each token's position is added, because otherwise the model could not tell 'dog bites man' from 'man bites dog'.

Step 3 - Transformer layers. The vectors pass through many layers. Each layer has an attention step (every token gathers information from relevant earlier tokens) and a feed-forward step (each token's vector is transformed independently, which is where much of the learned knowledge is thought to be stored). In a language model, attention is causal: a token can only look at tokens before it, never after, because at generation time the future does not exist yet.

Step 4 - Predict. The final vector of the last position is converted into one score per vocabulary token (called logits), and a function called softmax turns those scores into probabilities.

Step 5 - Sample. One token is chosen from the distribution (how exactly is covered in sampling-and-model-settings), appended to the input, and steps 1-5 repeat until the model produces a special end-of-turn token or hits the maximum length you allowed.

Training stages, summarised: pretraining (predict the next token on a vast text corpus - learns language and knowledge), instruction tuning (learn to follow instructions from example conversations), and preference training such as RLHF (learn which of several answers people prefer). The later stages are much smaller than pretraining but change the model's behaviour a lot.

Why it can be wrong: the model outputs the most plausible text given its training, and plausibility is not truth. Gaps or errors in training data, the training cutoff, ambiguous prompts and the randomness of sampling all lead to fluent but incorrect answers. Good AI engineering supplies the right facts in the prompt (RAG), lets the model use tools, and checks outputs.

"The cat sat on the"
        |  tokenise
        v
 [464][3797][3332][319][262]
        |  embed (+ position)
        v
 +-----------------------------+
 | transformer layer 1         |
 |  attention: look back       |
 |  feed-forward: transform    |
 +-----------------------------+
        |   ... many layers ...
        v
 scores for every token --> softmax
        v
 mat .41 | floor .18 | sofa .09 ...
        |  sample one
        v
 "mat"  --> append, repeat

Why does it exist?

Next-token prediction is a self-supervised task: the 'right answer' for every position is just the next token in existing text, so no human labelling is needed and you can train on almost unlimited data. Learning to predict text well forces the model to learn language, facts and reasoning patterns along the way.

The transformer exists because earlier sequence models processed text one word at a time and struggled to connect words far apart. Attention connects any two positions directly and can be computed in parallel on GPUs, which made training on huge datasets practical.

Instruction tuning and RLHF exist because a raw text predictor is not a helpful assistant; these stages align the model's behaviour with what users actually want.

When to use it

Understanding this mechanism helps you every day as an AI engineer: it explains why the prompt matters so much (it is the context for prediction), why output is billed per token, why long outputs take longer (one token at a time), why the model may be outdated (training cutoff), and why you must supply facts and verify outputs.

When not to use it

You do not need to understand the maths of transformers to build good applications; do not get stuck on matrix calculus before shipping your first prompt. Also avoid anthropomorphising: 'the model thinks/knows/wants' is convenient shorthand, but design your systems around what it actually is - a powerful next-token predictor shaped by training.

Common mistakes

  • Believing the model retrieves sentences from a stored copy of its training data; it generates from learned parameters.

  • Expecting the model to know recent events after its training cutoff without being told.

  • Assuming the model learns from your conversation permanently; parameters do not change during inference.

  • Thinking the model plans the whole answer before writing it - it produces one token at a time, so asking it to reason before answering can help.

  • Treating fluent, confident text as evidence of correctness.

  • Confusing a base model (continues text) with an instruction-tuned chat model (follows requests).

  • Assuming the model sees characters; it sees tokens, which is why letter-counting and spelling tasks can trip it up.

Practice exercises

  1. Easy:

    In the bigram demo, add two sentences to the corpus and print the new distribution after 'the'. Which probabilities changed and why?

  2. Easy:

    Explain to a friend, in four sentences, the difference between pretraining, instruction tuning and RLHF.

  3. Medium:

    Extend the bigram model into a trigram model: predict the next word from the previous TWO words. Does greedy generation still loop?

  4. Medium:

    In the attention demo, change the query vector so that 'it' attends most to 'street'. What does that correspond to in the sentence?

  5. Hard:

    Build a character-level bigram model: train it on a paragraph of text, then generate 80 characters by sampling (not greedy) using a seeded random number generator. Compare the output with the greedy version.

Interview questions

What is an LLM actually trained to do?

Predict the next token given the previous tokens. During pretraining it sees vast amounts of text and adjusts its parameters to make the real next token more likely. Generation repeats that prediction one token at a time.

What is attention, intuitively?

A mechanism that lets each token weigh how relevant every earlier token is and mix in information from the relevant ones. It is computed by comparing query and key vectors, normalising the scores with softmax, and averaging value vectors with those weights.

What is the difference between a base model and a chat/instruct model?

A base model has only been pretrained and continues text. A chat model has additionally been instruction-tuned and preference-trained (e.g. RLHF) to follow instructions, answer helpfully and refuse harmful requests.

What does RLHF do?

Reinforcement learning from human feedback: people compare model outputs, a reward signal is built from those preferences, and the model is optimised to produce outputs people prefer. It shapes helpfulness, tone, honesty and safety rather than teaching new knowledge.

Why do LLMs hallucinate?

They are trained to produce plausible continuations, not verified facts. With gaps or errors in training data, a training cutoff, ambiguous prompts or randomness in sampling, the most plausible text can be wrong, and the model has no built-in fact check.

Why is generation slower than reading the prompt?

The prompt's tokens can be processed in parallel, but output is autoregressive: each new token depends on the previous one, so tokens are produced one after another.

Does an LLM remember previous conversations?

Not by itself. Its parameters are fixed at inference time and the API is stateless. Any memory comes from the application sending previous messages or retrieved information again in the prompt.