Sampling, Temperature and Model Settings

How the next token is chosen (greedy, temperature, top-k, top-p), and the settings you actually control: max tokens, stop sequences, effort and thinking.

What is it?

At each step an LLM produces a probability distribution over possible next tokens (see how-llms-work). Sampling is the rule for picking one token from that distribution. The choice of rule controls whether output is predictable or varied.

  • Greedy decoding: always take the single most likely token. Predictable, but tends to be repetitive and bland.
  • Random sampling: pick a token at random, weighted by its probability. A token with probability 0.4 is picked 40% of the time. Varied, but occasionally picks something odd.
  • Temperature reshapes the distribution before sampling. The model's raw scores (logits) are divided by the temperature T before softmax. T < 1 sharpens the distribution (likely tokens become even more likely - closer to greedy). T > 1 flattens it (unlikely tokens get more chances - more surprising, more errors). T near 0 is effectively greedy.
  • Top-k: only consider the k most likely tokens, then sample among them.
  • Top-p (nucleus sampling): only consider the smallest set of top tokens whose probabilities add up to at least p (e.g. 0.9), then sample among them. This adapts: when the model is confident the set is tiny; when it is unsure the set is larger.

These knobs are general LLM concepts and many model APIs expose them as request parameters. However, some of the newest models fix their sampling internally and do not accept temperature or similar parameters - this includes the current Claude models used in this subject. Instead of tuning randomness, you steer them with the prompt and with an effort control.

Settings you will use with current Claude models:

  • max_tokens - the cap on generated tokens (including any thinking). Too low and replies get cut off (stop_reason == "max_tokens"); there is no charge for unused allowance, you pay for tokens actually produced.
  • stop_sequences - a list of strings; generation stops as soon as one is produced (stop_reason == "stop_sequence"). Useful for cutting output at a known delimiter.
  • system - the system prompt, which is the main way to control style and behaviour.
  • Effort - output_config={"effort": ...} with values low, medium, high, xhigh or max. Current models think adaptively (they may reason internally before answering); effort tells the model how much reasoning and thoroughness to apply. Lower effort means faster, cheaper, shorter work; higher effort means more careful reasoning on hard problems.

Determinism. Even at the lowest-randomness settings, LLM outputs are not guaranteed to be identical across calls (serving infrastructure, batching and model updates all play a part). Design systems that tolerate variation: validate outputs, use structured outputs for machine-readable data, and test with several runs.

Practical intuition for models that do expose temperature: low values for extraction, classification and factual Q&A; moderate values for general chat; higher values for brainstorming and creative writing - and change one setting at a time while measuring on a test set.

Explain like I'm 10

Picture the model holding a bag of marbles for the next word, with more marbles for likelier words. Greedy picks the colour with the most marbles every time. Temperature is like adding or removing marbles: low temperature gives the favourite colour almost all the marbles; high temperature evens the bag out. Top-k says 'throw out every colour except the k biggest piles'; top-p says 'keep only the biggest piles that together make up 90% of the bag'. Effort is different: it is like telling the player how long to think before drawing at all.

Examples

Temperature sampling over a fixed distribution (seeded)

// Seeded PRNG so results are reproducible
function mulberry32(seed) {
  return function () {
    seed |= 0; seed = (seed + 0x6D2B79F5) | 0;
    let t = Math.imul(seed ^ (seed >>> 15), 1 | seed);
    t = (t + Math.imul(t ^ (t >>> 7), 61 | t)) ^ t;
    return ((t ^ (t >>> 14)) >>> 0) / 4294967296;
  };
}

// Pretend logits (raw scores) for the word after "The weather today is"
const tokens = ["sunny", "cloudy", "rainy", "lovely", "purple"];
const logits = [3.0, 2.2, 1.8, 1.0, -1.0];

function softmax(scores, temperature) {
  const scaled = scores.map(s => s / temperature);
  const max = Math.max(...scaled);              // subtract max for numerical stability
  const exps = scaled.map(s => Math.exp(s - max));
  const sum = exps.reduce((a, b) => a + b, 0);
  return exps.map(e => e / sum);
}

function sample(probs, rand) {
  let r = rand(), cumulative = 0;
  for (let i = 0; i < probs.length; i++) {
    cumulative += probs[i];
    if (r < cumulative) return i;
  }
  return probs.length - 1;
}

for (const T of [0.2, 1.0, 2.0]) {
  const probs = softmax(logits, T);
  const rand = mulberry32(7);
  const counts = new Array(tokens.length).fill(0);
  for (let i = 0; i < 1000; i++) counts[sample(probs, rand)]++;
  console.log("T=" + T);
  tokens.forEach((t, i) =>
    console.log("  " + t.padEnd(7), "p=" + probs[i].toFixed(3), "picked", counts[i], "/1000"));
}

At T=0.2 'sunny' wins almost every time (near-greedy). At T=2.0 even 'purple' gets picked sometimes - more variety, more nonsense. That is the whole trade-off temperature controls.

Top-k and top-p filtering

const dist = [
  ["sunny", 0.45], ["cloudy", 0.25], ["rainy", 0.15],
  ["lovely", 0.08], ["windy", 0.05], ["purple", 0.02],
];

function topK(d, k) {
  const kept = [...d].sort((a, b) => b[1] - a[1]).slice(0, k);
  const total = kept.reduce((s, x) => s + x[1], 0);
  return kept.map(([t, p]) => [t, +(p / total).toFixed(3)]); // renormalise
}

function topP(d, p) {
  const sorted = [...d].sort((a, b) => b[1] - a[1]);
  const kept = [];
  let cumulative = 0;
  for (const item of sorted) {
    kept.push(item);
    cumulative += item[1];
    if (cumulative >= p) break;
  }
  const total = kept.reduce((s, x) => s + x[1], 0);
  return kept.map(([t, q]) => [t, +(q / total).toFixed(3)]);
}

console.log("top-k=2  :", JSON.stringify(topK(dist, 2)));
console.log("top-p=0.8:", JSON.stringify(topP(dist, 0.8)));
console.log("top-p=0.95:", JSON.stringify(topP(dist, 0.95)));

// A confident distribution: top-p keeps just one token
const confident = [["Paris", 0.97], ["Lyon", 0.02], ["Rome", 0.01]];
console.log("confident, top-p=0.9:", JSON.stringify(topP(confident, 0.9)));

Top-k always keeps k tokens. Top-p adapts to the model's confidence: a confident distribution keeps one token, an uncertain one keeps several. Either way, the long tail of unlikely (often wrong) tokens is cut off.

Settings you control on current Claude models (Python)

import anthropic

client = anthropic.Anthropic()

# 1) Quick, cheap task: low effort
quick = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=16000,
    output_config={"effort": "low"},
    messages=[{"role": "user", "content": "Give me a title for a blog post about unit testing."}],
)

# 2) Hard reasoning task: high effort (more internal thinking)
careful = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=16000,
    output_config={"effort": "high"},
    messages=[{"role": "user", "content": "Find the bug in this binary search: ..."}],
)

# 3) Stop sequences: cut the output at a delimiter
listing = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=16000,
    stop_sequences=["END_OF_LIST"],
    messages=[{"role": "user", "content": "List three fruits, one per line, then write END_OF_LIST."}],
)
print(listing.stop_reason)   # "stop_sequence" if the delimiter was produced

for r in (quick, careful, listing):
    print(r.stop_reason, r.usage.output_tokens)

On current Claude models you do not set temperature; you control cost and depth with effort, length with max_tokens, and cut-off points with stop_sequences. Compare output token counts between low and high effort.

How it works

The model's last layer outputs one logit (raw score) per vocabulary token. The decoding step then: (1) divides logits by the temperature, (2) applies softmax to get probabilities, (3) optionally removes tokens outside the top-k or top-p set and renormalises, (4) draws one token at random according to the remaining probabilities. Greedy decoding skips the randomness and takes the maximum.

Why not always use greedy? Always picking the top token often produces repetitive, dull or looping text (you saw this in the bigram demo), and it can lock the model into a poor early choice. A little randomness tends to produce more natural text.

Why do some newest models remove temperature? Their training and serving are tuned for a particular decoding setup, and their behaviour is better steered by instructions and an effort control than by low-level randomness knobs. The concepts still matter: they explain why the same prompt can give different answers, and you will see these parameters on many other models and local inference tools.

Thinking and effort. Current models can generate internal reasoning before the visible answer. With adaptive thinking, the model decides how much to think; the effort setting biases it toward less (fast, cheap) or more (thorough). Thinking tokens count as output tokens toward max_tokens and cost, so give enough headroom when using high effort.

logits:  sunny 3.0  cloudy 2.2  rainy 1.8  purple -1.0
             |  divide by temperature T
             v
softmax:  T=0.2 -> sunny .98 ...   (sharp, ~greedy)
          T=1.0 -> sunny .52 ...   (as trained)
          T=2.0 -> sunny .37 ...   (flat, wild)
             |  top-k / top-p: cut the long tail
             v
          random pick -> "sunny"
             |
          append, repeat until end / max_tokens / stop seq

Why does it exist?

A model only gives probabilities; something has to turn them into actual text. Sampling strategies exist to balance quality (pick likely tokens) against diversity (do not always pick the same one). Output settings like max_tokens and stop sequences exist to control cost and shape, and effort exists because different tasks deserve different amounts of reasoning.

When to use it

Use low effort for simple, high-volume tasks (classification, short rewrites, routing) and higher effort for complex reasoning, coding and agentic work. Use stop sequences to end output at a known delimiter. On models that expose temperature, lower it for extraction and factual tasks and raise it for brainstorming; measure the effect on a test set.

When not to use it

Do not rely on low randomness as a correctness guarantee - a wrong answer can be the most likely one. Do not pass temperature, top_p or top_k to models that do not support them. Do not set max_tokens as a way to request brevity (it truncates mid-sentence); ask for a short answer in the prompt instead.

Common mistakes

  • Expecting identical outputs for identical prompts; build validation and tests that tolerate variation.

  • Using max_tokens to make answers short, producing cut-off text instead of concise text.

  • Setting max_tokens too low with high effort, so thinking uses the budget and the answer is truncated.

  • Tuning temperature and top-p at the same time and not knowing which change helped.

  • Sending sampling parameters a model no longer supports and getting request errors.

  • Using maximum effort for every trivial request and paying for unnecessary reasoning.

Practice exercises

  1. Easy:

    Run the temperature demo with T = 0.5 and T = 5. Describe how the counts change and what kind of text each would produce.

  2. Easy:

    Explain the difference between top-k and top-p using the 'confident' distribution in the demo.

  3. Medium:

    Combine the demos: write sampleToken(logits, {temperature, topK, topP}, rand) that applies temperature, then top-k, then top-p, then samples. Test it with the seeded PRNG.

  4. Medium:

    Call the API with the same prompt at effort low, medium and high on a tricky logic puzzle. Record correctness, output tokens and latency for each.

  5. Hard:

    Build a toy text generator: train a word bigram model on a few paragraphs, then generate 30 words with your sampleToken function at several temperatures. Show how repetition and nonsense change with temperature.

Interview questions

What does temperature do?

It divides the logits before softmax. Below 1 it sharpens the distribution toward the most likely tokens (more deterministic); above 1 it flattens it (more diverse and more error-prone). Near 0 it approaches greedy decoding.

Explain top-p (nucleus) sampling.

Sort tokens by probability and keep the smallest set whose cumulative probability reaches p, renormalise and sample from that set. It adapts to confidence: few candidates when the model is sure, more when unsure.

Does setting temperature to 0 make outputs fully deterministic and correct?

No. It makes decoding close to greedy, but outputs can still vary across runs due to infrastructure, and the most likely answer can be wrong. Correctness needs grounding and validation.

How do you control reasoning depth on current Claude models?

Thinking is adaptive and on by default; you steer it with the effort setting via output_config={"effort": ...} (low to max). Higher effort spends more tokens on reasoning; make sure max_tokens leaves enough room.

What is a stop sequence used for?

To end generation as soon as a specific string is produced, for example a delimiter after a list. The response's stop_reason becomes stop_sequence.

Why not always use greedy decoding?

Greedy output tends to be repetitive and can get stuck in loops or a bad early choice. Controlled randomness generally produces more natural and diverse text.