Workflow Patterns

Five proven ways to compose LLM calls in code: prompt chaining, routing, parallelization, orchestrator-workers and evaluator-optimizer.

What is it?

Most successful LLM features are not free-roaming agents; they are workflows - LLM calls composed by ordinary code in a known structure. Workflows are cheaper, faster, easier to test and easier to debug than agents, and they cover a large share of real use cases. Five patterns come up again and again. Learn them, and you will recognise which one a problem needs.

All five are built from the same unit, sometimes called the augmented LLM: a model call that may also have retrieval, tools and memory. The patterns differ only in how your code connects those calls.

1. Prompt chaining. Split a task into fixed sequential steps, where each call works on the output of the previous one: outline, then check the outline, then write; or extract, then translate, then format. Between steps you can add a gate - a plain code check (length, valid JSON, required sections) that stops or retries before errors propagate. Each step is a simpler task, so each is more accurate. Trade-off: more latency.

2. Routing. Classify the input first, then send it to a specialised handler: refunds to a billing prompt, bugs to a technical prompt with docs search, small talk to a cheap fast model. Routing lets each path have its own prompt, tools and model size without one prompt trying to do everything. The classifier can be an LLM call or a traditional classifier.

3. Parallelization. Run independent LLM calls at the same time and combine them in code. Two flavours:

  • Sectioning: split a task into independent parts handled in parallel - for example one call answers the user while another screens the request against policy (a guardrail), or each section of a long document is reviewed separately.
  • Voting: run the same task several times (or with different prompts) and aggregate - majority vote, or 'flag if any reviewer flags'. This trades cost for confidence and lets you tune false positives vs false negatives.

4. Orchestrator-workers. A central LLM (the orchestrator) looks at the task, decides dynamically what subtasks are needed, hands each to a worker LLM call, and synthesises the results. Unlike sectioning, the subtasks are not known in advance - for example, which files need to change for a feature request depends on the request. It sits between workflow and agent: the plan is dynamic, but the structure (plan, fan out, combine) is fixed.

5. Evaluator-optimizer. One call generates, another evaluates against clear criteria and returns specific feedback, and the generator revises - looping until the evaluator passes it or a round limit is hit. It works when you can articulate what 'good' means and when feedback measurably improves output: translations, code that must pass tests, writing that must meet a checklist.

These patterns combine: a router whose technical branch uses a chain whose last step is an evaluator-optimizer loop. Start with the simplest pattern that solves the problem, measure, and add structure only where measurements show it helps. In production, the 'parallel' calls really run concurrently - Promise.all in JavaScript or asyncio.gather / a thread pool in Python; the runnable sketches below call a fake llm() sequentially to stay simple and deterministic.

Explain like I'm 10

A restaurant kitchen. Chaining is the line: prep, cook, plate, with the head chef checking each plate (the gate). Routing is the host sending diners to the bar or the dining room. Parallelization is several cooks working different dishes at once, or three tasters voting on a new sauce. Orchestrator-workers is the head chef reading a big catering order and deciding who cooks what. Evaluator-optimizer is a chef and a critic going back and forth until the dish is right.

Examples

Prompt chaining with a gate, and routing (runnable)

// Fake LLM with canned behaviour per instruction prefix.
function llm(prompt) {
  if (prompt.startsWith("OUTLINE:")) return prompt.includes("caching")
    ? "What caching is; Cache invalidation; When not to cache"
    : "Intro";                                         // a weak outline for other topics
  if (prompt.startsWith("WRITE:")) return "Article with sections -> " + prompt.slice("WRITE: ".length);
  if (prompt.startsWith("ROUTE:")) {
    const t = prompt.toLowerCase();
    if (t.includes("refund") || t.includes("charged")) return "billing";
    if (t.includes("error") || t.includes("crash")) return "technical";
    return "general";
  }
  return "";
}

// ---- 1. PROMPT CHAINING: outline -> gate (plain code) -> write ----
function writeArticle(topic) {
  const outline = llm("OUTLINE: " + topic);
  const sections = outline.split(";").length;
  if (sections < 3) return "GATE FAILED for '" + topic + "': only " + sections + " section(s); not writing";
  return llm("WRITE: " + outline);
}
console.log(writeArticle("HTTP caching"));
console.log(writeArticle("DNS"));

// ---- 2. ROUTING: classify, then hand to a specialised handler ----
const handlers = {
  billing: (t) => "billing prompt + refund tools, small fast model",
  technical: (t) => "tech prompt + docs search, larger model",
  general: (t) => "FAQ prompt, small fast model",
};
for (const ticket of ["I was charged twice", "App shows error 500 on login", "Do you ship to Canada?"]) {
  const route = llm("ROUTE: " + ticket);
  console.log(route.padEnd(9) + " <- " + ticket + "  => " + handlers[route](ticket));
}

The chain's gate is ordinary code, so a bad intermediate result stops the pipeline instead of producing a bad article. The router means each handler can be small and focused.

Parallelization: sectioning and voting (runnable)

// Fake LLM: behaviour depends on the instruction.
function llm(prompt) {
  if (prompt.startsWith("ANSWER:")) return "Here is how to reset your password: ...";
  if (prompt.startsWith("SCREEN:")) return prompt.includes("password of another user") ? "UNSAFE" : "SAFE";
  if (prompt.startsWith("REVIEW-SQL:")) return prompt.includes("' + userId") ? "VULNERABLE" : "OK";
  if (prompt.startsWith("REVIEW-XSS:")) return prompt.includes("innerHTML") ? "VULNERABLE" : "OK";
  if (prompt.startsWith("REVIEW-GENERAL:")) return prompt.includes("+ userId") ? "VULNERABLE" : "OK";
  return "";
}
// In production these run concurrently: await Promise.all([...]) or asyncio.gather(...).
const parallel = (prompts) => prompts.map(llm);

// ---- SECTIONING: answer and guardrail screening at the same time ----
for (const q of ["How do I reset my password?", "How do I get the password of another user?"]) {
  const [answer, verdict] = parallel(["ANSWER: " + q, "SCREEN: " + q]);
  console.log(q, "->", verdict === "SAFE" ? answer : "Sorry, I can't help with that.");
}

// ---- VOTING: three differently-focused reviewers, aggregate in code ----
const snippet = "db.query('SELECT * FROM users WHERE id=' + userId)";
const votes = parallel(["REVIEW-SQL: ", "REVIEW-XSS: ", "REVIEW-GENERAL: "].map((p) => p + snippet));
const flagged = votes.filter((v) => v === "VULNERABLE").length;
console.log("votes:", votes.join(", "));
console.log("majority rule  :", flagged >= 2 ? "FLAG" : "pass");
console.log("any-flag rule  :", flagged >= 1 ? "FLAG" : "pass", "(more sensitive, more false alarms)");

Sectioning splits different jobs across parallel calls; voting runs the same question through several reviewers and lets your code choose the aggregation rule - majority for precision, any-flag for recall.

Orchestrator-workers: a dynamic plan, fixed structure (runnable)

const codebase = {
  "settings.ts": "user settings model",
  "SettingsPage.tsx": "settings screen UI",
  "theme.css": "colour variables",
  "api/users.ts": "user REST endpoints",
};

// ORCHESTRATOR (fake LLM): decides which subtasks this request needs.
function orchestrator(request) {
  const r = request.toLowerCase();
  const plan = [];
  if (r.includes("setting")) plan.push({ file: "settings.ts", job: "add the new field with a default" },
                                       { file: "SettingsPage.tsx", job: "add a control for it" });
  if (r.includes("dark") || r.includes("theme")) plan.push({ file: "theme.css", job: "add dark colour variables" });
  if (r.includes("api") || r.includes("sync")) plan.push({ file: "api/users.ts", job: "expose the field in the API" });
  return plan;
}
// WORKER (fake LLM): sees only its own file and job - a small, focused context.
function worker(sub) {
  return sub.file + ": " + sub.job + " [context: " + codebase[sub.file] + "]";
}
// SYNTHESIZER (fake LLM): combines worker outputs into one result.
function synthesize(request, outputs) {
  return "Change plan for '" + request + "' (" + outputs.length + " workers):\n  - " + outputs.join("\n  - ");
}

for (const request of ["Add a dark mode setting", "Sync the language setting via the API"]) {
  const plan = orchestrator(request);
  const outputs = plan.map(worker);          // parallel in production
  console.log(synthesize(request, outputs));
}

The two requests produce different numbers and kinds of subtasks - that is what distinguishes orchestrator-workers from fixed sectioning. In a real system the orchestrator returns its plan as structured output (a JSON list of subtasks).

Evaluator-optimizer: generate, critique, revise (runnable)

// GENERATOR (fake LLM): improves the draft using all feedback so far.
function generator(task, feedback) {
  let draft = "slugify(s) lowercases s and replaces spaces with dashes.";
  if (feedback.includes("example")) draft += " Example: slugify('Hello World') returns 'hello-world'.";
  if (feedback.includes("edge cases")) draft += " Edge cases: trims spaces; drops characters other than a-z, 0-9 and dashes.";
  return draft;
}
// EVALUATOR (fake LLM): checks explicit criteria, returns the first problem found.
function evaluator(draft) {
  if (!draft.includes("Example")) return { pass: false, feedback: "add an example" };
  if (!draft.includes("Edge cases")) return { pass: false, feedback: "mention edge cases" };
  return { pass: true, feedback: "" };
}

const task = "Document the slugify function";
let allFeedback = [];
const MAX_ROUNDS = 4;
for (let round = 1; round <= MAX_ROUNDS; round++) {
  const draft = generator(task, allFeedback.join("; "));
  const verdict = evaluator(draft);
  console.log("round " + round + ": " + (verdict.pass ? "PASS" : "revise -> " + verdict.feedback));
  if (verdict.pass) { console.log("final:", draft); break; }
  allFeedback.push(verdict.feedback);
  if (round === MAX_ROUNDS) console.log("gave up after", MAX_ROUNDS, "rounds; returning best draft");
}

The loop has a clear exit (the evaluator passes) and a hard cap (MAX_ROUNDS). The evaluator's criteria must be explicit; 'make it better' never converges.

How it works

Every pattern is a few lines of control flow around llm() calls: sequential calls with checks (chaining), a classification followed by a dispatch table (routing), concurrent calls plus an aggregation function (parallelization), a planning call that returns a list followed by a fan-out and a synthesis call (orchestrator-workers), and a bounded loop of generate and evaluate (evaluator-optimizer).

Because the structure lives in your code, each step can be tested and measured on its own: does the router classify correctly, does the gate catch bad outlines, how often does the evaluator pass on round one? That observability is the main advantage over a free-form agent.

To make intermediate results reliable, have steps return structured output (a schema for the classification, a JSON plan for the orchestrator, a pass/fail plus feedback object for the evaluator), and validate before the next step.

CHAINING      in -> [LLM] -> gate -> [LLM] -> out
ROUTING       in -> [classify] -> A | B | C -> out
SECTIONING    in -> [LLM a] \
                 -> [LLM b]  > combine -> out
VOTING        in -> [LLM]x3 -> majority -> out
ORCH-WORKERS  in -> [plan] -> [w1][w2]..[wn] -> [merge]
EVAL-OPT      in -> [gen] -> [eval] -pass-> out
                      ^--feedback--+ (max rounds)

Why does it exist?

A single prompt asked to do everything at once is harder for the model and impossible to debug. A free agent is flexible but costly and unpredictable. Workflow patterns sit in between: they decompose problems into steps the model handles well, while keeping the control flow explicit, testable and cheap.

When to use it

Chaining: tasks with clear fixed stages. Routing: distinct input categories that deserve different handling or model sizes. Sectioning: independent subtasks or a guardrail alongside the main answer. Voting: high-stakes judgments where you want confidence. Orchestrator-workers: complex tasks whose subtasks depend on the input. Evaluator-optimizer: outputs with articulable quality criteria that improve with feedback.

When not to use it

Do not add a pattern a single prompt already handles well - each extra call adds latency and cost. Avoid voting when calls are expensive and the task is low-stakes. Avoid evaluator-optimizer when you cannot state concrete criteria. If the number and kind of steps are truly unpredictable and need environment feedback, move to an agent.

Common mistakes

  • Chaining steps without validation gates, so an early error flows silently to the end.

  • Routing with a vague classifier prompt and no 'other' category, so odd inputs get forced into the wrong path.

  • Running independent calls sequentially and paying the latency of each one in turn.

  • Voting with identical prompts and identical settings, which often produces near-identical answers and little extra signal.

  • Evaluator-optimizer loops without a maximum number of rounds, or with vague criteria that never converge.

  • Letting the orchestrator return free text instead of a structured plan that code can validate.

  • Building an agent when one of these fixed patterns would be cheaper and more reliable.

Practice exercises

  1. Easy:

    Classify each as chaining, routing, sectioning, voting, orchestrator-workers or evaluator-optimizer: a translator with a reviewer loop; a support bot that sends billing questions to one prompt and tech questions to another; generating marketing copy then translating it.

  2. Easy:

    In the routing demo, add a 'sales' route for messages containing 'price' or 'quote' and test it.

  3. Medium:

    Rewrite the parallelization demo with async functions and Promise.all, where each fake llm() call waits with setTimeout, and measure total time vs running them one after another.

  4. Medium:

    Implement a real prompt chain in Python with the Anthropic SDK: extract action items from meeting notes as structured output, gate on at least one item, then draft a follow-up email.

  5. Hard:

    Build a real evaluator-optimizer for SQL generation: the evaluator runs the query against a SQLite test database and returns errors or wrong row counts as feedback. Cap at 4 rounds and report how often each round succeeds.

  6. Hard:

    Combine patterns: route incoming questions; send simple ones to one-shot RAG and complex ones to orchestrator-workers that research sub-questions in parallel. Measure latency and cost per path.

Interview questions

Name and describe the five common LLM workflow patterns.

Prompt chaining (fixed sequential steps with gates), routing (classify then dispatch to specialised handlers), parallelization (sectioning independent subtasks or voting on the same task), orchestrator-workers (an LLM dynamically plans subtasks, workers execute, results are synthesised), and evaluator-optimizer (generate, evaluate against criteria, revise in a bounded loop).

How is orchestrator-workers different from sectioning?

In sectioning the subtasks are predefined by the developer. In orchestrator-workers an LLM decides at run time which subtasks are needed based on the input, so the number and nature of workers varies.

When does evaluator-optimizer work well?

When there are clear, checkable criteria and feedback demonstrably improves results - for example code that must pass tests, translations with specific quality rubrics, or text that must meet a checklist. It needs a maximum round count.

Why prefer a workflow over an agent?

Workflows are predictable, cheaper, lower latency, and each step can be tested and monitored. Agents are only worth their cost when the steps cannot be determined in advance.

What is a gate in prompt chaining?

A programmatic check between steps - schema validation, length, required fields, a classifier - that stops, retries or branches before a bad intermediate result propagates to later steps.

How do you aggregate votes?

Choose the rule to match the cost of errors: majority vote for balanced precision, any-flag for high recall on safety issues, all-agree for high precision, or a threshold. Diversify prompts or perspectives so voters are not making identical errors.