Prompt Engineering Basics
How to write prompts that work: system vs user messages, clear instructions, context, examples (few-shot), XML-tag structure and asking for step-by-step reasoning.
What is it?
A prompt is everything you send the model: instructions, context, examples, and the actual question. Since an LLM predicts the most plausible continuation of its input (see how-llms-work), the prompt is the program. Prompt engineering is the practice of writing that input so the model reliably does what you want.
Chat APIs split the prompt into roles:
- System prompt: standing instructions that apply to the whole conversation - the model's role, rules, tone, output format, what to do when unsure. Set by you, the developer, not by the end user.
- User messages: the requests (from your end user or from your code).
- Assistant messages: the model's previous replies, which you send back to give it the conversation history.
The single most useful rule: be clear and specific, as if briefing a smart new colleague who knows nothing about your project. A vague prompt ('summarise this') gets a generic result. A specific prompt says who the output is for, what to include, how long, what format, and what to do in edge cases.
Techniques that consistently help:
- Give context and the reason: 'Summarise this incident report for the on-call engineer who needs to decide whether to page the database team' beats 'summarise'. Knowing why lets the model make good judgement calls.
- Say what to do, not only what not to do: 'Write in plain prose paragraphs' works better than 'don't use bullet points'.
- Show examples (few-shot prompting): 2-5 examples of input and ideal output teach format and style far more precisely than description alone. Make them varied so the model does not copy one example too literally.
- Structure with XML tags: wrap different parts in labelled tags such as
<document>,<instructions>,<example>. This prevents the model from confusing your instructions with the data it should process, and lets you refer to parts by name ('using the report in the <report> tags'). - Put long documents first, the question last: with lots of material, place the documents at the top and your instructions and question at the end.
- Ask for reasoning before the answer when the task needs thinking: 'First work through the problem step by step, then give the final answer.' This is often called chain-of-thought. Many current models also have built-in thinking that you control with an effort setting (see sampling-and-model-settings).
- Define the output format precisely: give a template or a schema; for machine-readable output use structured outputs (see structured-output).
- Allow uncertainty: tell the model it may say it does not know, or that information is missing, instead of guessing.
Prompting is iterative engineering, not magic words. Write a prompt, run it on a set of realistic test inputs (including awkward ones), look at the failures, adjust, and run again. Keep prompts in version control and keep the test inputs, because a change that fixes one case can break another.
A common trap is the 'prompt that only works on the demo'. Real users paste messy text, ask off-topic questions, and write in other languages. Test with those.
Explain like I'm 10
Prompting is like delegating to a brilliant contractor who has never visited your company. If you say 'fix the website', you get something - but probably not what you wanted. If you say 'the checkout button on mobile overlaps the price; customers are abandoning carts; keep the brand colours; here are two pages that look right', you get exactly what you need. The system prompt is the contractor's standing brief; examples are 'here is what good looks like'; XML tags are labelled folders so nothing gets mixed up.
Examples
Assembling a well-structured prompt (offline)
// Build a prompt from parts. The output is what you would send to the model.
function buildPrompt({ document, question, examples }) {
const shots = examples
.map(e => "<example>" + " input: " + e.input + " | output: " + e.output + " </example>")
.join(" ");
return [
"<document>",
document,
"</document>",
"",
"<examples>",
shots,
"</examples>",
"",
"<instructions>",
"Answer the question using only the document above.",
"If the document does not contain the answer, reply exactly: NOT IN DOCUMENT.",
"Answer in one sentence, in the style of the examples.",
"</instructions>",
"",
"Question: " + question,
].join(" ~ "); // using " ~ " as a visible line separator for this demo
}
const prompt = buildPrompt({
document: "Acme Cloud refunds are available within 30 days of purchase. Annual plans are refunded pro rata.",
question: "Can I get a refund after 2 weeks?",
examples: [
{ input: "Is there a free tier?", output: "NOT IN DOCUMENT." },
{ input: "How long is the refund window?", output: "Refunds are available within 30 days of purchase." },
],
});
for (const line of prompt.split(" ~ ")) console.log(line);
console.log("---");
console.log("approx tokens:", Math.ceil(prompt.length / 4));Notice the structure: data in tags at the top, examples showing both a normal answer and the 'not found' answer, explicit instructions including the fallback, and the question last. This shape is reused in almost every RAG prompt later in the subject.
System prompt, XML tags and few-shot examples (Python)
import anthropic
client = anthropic.Anthropic()
SYSTEM = """You classify customer support tickets for an online bookshop.
The labels go straight into our routing system, so reply with exactly one
label from this list and nothing else: shipping, billing, account, other.
If a ticket mentions several issues, choose the one the customer is most upset about."""
# Few-shot examples are sent as earlier conversation turns.
examples = [
{"role": "user", "content": "<ticket>My parcel says delivered but it is not here.</ticket>"},
{"role": "assistant", "content": "shipping"},
{"role": "user", "content": "<ticket>I was charged twice for the same order!!</ticket>"},
{"role": "assistant", "content": "billing"},
]
ticket = "I can't log in since I changed my email, and also the book arrived late."
resp = client.messages.create(
model="claude-opus-5-5",
max_tokens=16000,
system=SYSTEM,
messages=examples + [{"role": "user", "content": f"<ticket>{ticket}</ticket>"}],
)
label = next(b.text for b in resp.content if b.type == "text").strip()
print(label) # most likely "account"The system prompt explains the task AND why the format matters (it feeds a routing system), gives the allowed labels, and resolves the ambiguous case. The examples demonstrate the exact output shape. In production, also validate the label in code (see structured-output).
Asking for reasoning before the answer (Python)
import anthropic
client = anthropic.Anthropic()
prompt = """<data>
Plan A: $12 per month, 3 users included, $5 per extra user.
Plan B: $30 per month, 10 users included, $2 per extra user.
</data>
Our team has 8 users today and expects 14 users next year.
Which plan is cheaper today, and which is cheaper next year?
Work through the calculation step by step inside <reasoning> tags,
then give the final recommendation in <answer> tags."""
resp = client.messages.create(
model="claude-opus-5-5",
max_tokens=16000,
messages=[{"role": "user", "content": prompt}],
)
text = next(b.text for b in resp.content if b.type == "text")
# Show only the answer part to the end user
start, end = text.find("<answer>"), text.find("</answer>")
print(text[start + len("<answer>"):end].strip() if start != -1 else text)Writing out the steps gives the model 'room to think' in tokens before committing to an answer, which helps on multi-step problems. Separating reasoning and answer with tags lets your code show only the final part.
How it works
The model conditions every generated token on the entire prompt. Clear instructions and examples shift the probabilities toward the outputs you want; ambiguity spreads probability over many reasonable-but-different outputs, which you experience as inconsistency.
System prompts are given special weight by chat-trained models because instruction tuning taught them that the system role holds the developer's rules. Examples work because the model is extremely good at continuing patterns (few-shot learning emerged from pretraining on text full of patterns). XML tags work because models were trained on lots of structured text and learn to treat tagged sections as distinct units.
Step-by-step reasoning helps because each generated token is computed with a fixed amount of work. Producing intermediate steps lets the model spread a hard problem across many tokens and then condition the final answer on its own worked steps.
A practical workflow: (1) write the task as if briefing a colleague; (2) collect 10-30 realistic test inputs including edge cases; (3) run and read the outputs; (4) fix the prompt where it failed, preferring explanations over shouting ('IMPORTANT!!!' tends to cause over-application); (5) re-run everything to catch regressions.
+-------------------- prompt --------------------+
| SYSTEM: role, rules, format, why, fallbacks |
+------------------------------------------------+
| USER: <document> long material </document> |
| <examples> input -> output </examples> |
| <instructions> what to do </instructions>|
| Question goes LAST |
+------------------------------------------------+
|
v
model predicts reply
|
check output -> adjust prompt -> rerunWhy does it exist?
The same model can produce a mediocre or an excellent result depending only on its input. Prompt engineering exists because it is the cheapest, fastest lever you have: no training, no infrastructure, changes deploy instantly. Most tasks that people assume need fine-tuning are solved by a well-written prompt plus the right context.
When to use it
Always - every LLM feature starts with a prompt. Invest more effort when outputs feed other systems (strict format), when users are external (robustness to odd inputs), and when mistakes are expensive. Use few-shot examples whenever the exact output style matters; use XML tags whenever the prompt mixes instructions with data; ask for reasoning on multi-step or mathematical tasks.
When not to use it
Do not rely on prompting alone to enforce hard guarantees (valid JSON, allowed values, security rules) - validate in code and use structured outputs and permissions. Do not try to fix missing knowledge with clever wording; supply the knowledge with retrieval (what-is-rag). Do not chase 'magic phrases' from the internet instead of testing on your own data.
Common mistakes
Writing vague instructions ('make it better') and blaming the model for generic results.
Mixing instructions and user-supplied data without separating them, so text inside the data gets treated as instructions.
Using a single example, which the model copies too literally; use several varied examples.
Shouting with CAPITALS and 'NEVER' everywhere, which can make the model over-cautious or rigid; explain the reason instead.
Testing a prompt on one or two happy-path inputs only.
Putting the question before a long document instead of after it.
Not telling the model what to do when information is missing, so it guesses.
Changing prompts in production without re-running a test set.
Practice exercises
- Easy:
Rewrite the prompt 'Summarise this email' into a specific prompt that names the audience, length, format and what to do if the email contains no action items.
- Easy:
Take the offline prompt builder and add a <context> section describing the reader (a customer on a mobile phone). Print the result.
- Medium:
Write a system prompt and three few-shot examples for extracting the sender's name, company and request from an email, in a fixed three-line format. List five tricky test emails to try it on.
- Medium:
Build a JS function
extractTag(text, tag)that returns the content between <tag> and </tag>, or null if missing. Test it on text with zero, one and two answer blocks. - Hard:
Build a tiny prompt test harness in Python: a list of (input, expected_label) pairs, a function that calls the API with your classification prompt, and a report of accuracy plus every failing case. Improve the prompt until accuracy stops improving.
Interview questions
What is the difference between the system prompt and a user message?
The system prompt holds the developer's standing instructions for the whole conversation (role, rules, format, policies). User messages are individual requests. Models are trained to treat the system prompt as higher-priority context.
What is few-shot prompting and when do you use it?
Including several input/output examples in the prompt so the model infers the pattern. Use it when format, style or labelling conventions matter. Use varied examples, including edge cases, to avoid the model copying one example too closely.
Why structure prompts with XML tags?
Tags clearly separate instructions, documents, examples and inputs, which reduces confusion and makes it harder for data to be mistaken for instructions. They also make it easy to reference sections and to parse tagged output.
Why can asking the model to reason step by step improve answers?
Each token gets a bounded amount of computation. Writing intermediate steps lets the model spread the work across many tokens and base the final answer on its own explicit reasoning, which helps on multi-step problems.
How do you know a prompt change is an improvement?
Run it against a fixed set of representative test inputs with known good outputs, compare metrics (accuracy, format validity) before and after, and inspect failures. One good-looking example proves little.
How do you reduce made-up answers through prompting?
Provide the relevant source material, instruct the model to answer only from it, explicitly allow 'I don't know' with an exact fallback phrase, and ask for supporting quotes. Then verify in code where possible.