Structured Output: JSON, Schemas and Validation

Getting machine-readable output from an LLM: JSON, JSON Schema, Pydantic models, the SDK's parse helper, validation and retry.

What is it?

Most LLM output in real applications is read by code, not people: a ticket category to route, fields extracted from an invoice, a list of tags to save in a database. Code needs structured output - data in a predictable shape - not a friendly paragraph.

The usual format is JSON (JavaScript Object Notation): text that represents objects {"key": value}, arrays [1, 2, 3], strings, numbers, booleans and null. Almost every language can parse it.

A schema describes the shape the JSON must have: which fields exist, their types, which are required, allowed values. JSON Schema is the standard language for that, and tools like Pydantic (Python) or Zod (TypeScript) let you define the shape as a class or object and validate data against it.

There are three levels of reliability, from weakest to strongest:

  • Ask nicely: 'Reply with JSON containing category and urgent.' Usually works, but you may get extra prose, markdown code fences, missing fields, or a different spelling of a value.
  • Ask with a schema and examples, then validate and retry: include the schema in the prompt, parse the result, validate it, and on failure send the error back and ask again. Works with any model.
  • Use the API's structured-output feature: you pass the schema with the request and the API constrains the model's output to match it, so the response parses and conforms to the schema's structure. With the Anthropic Python SDK, client.messages.parse(..., output_format=YourPydanticModel) does this and returns the parsed object in resp.parsed_output.

Even with guaranteed structure, validate the meaning in code. A schema can guarantee amount is a number; it cannot guarantee the number is the right one from the invoice. Business rules (amount > 0, due date after issue date, the category exists in your system) are your code's job.

Design tips for schemas: keep them as flat and small as the task allows; use enums (a fixed list of allowed values) for categories; give fields clear names and descriptions; allow null for information that may genuinely be missing, so the model does not have to invent a value; and if you want the model to think before filling fields, add a reasoning field before the answer fields.

Explain like I'm 10

Asking for free-text answers is like asking a colleague to 'tell me about the invoice' over the phone - you get a story you have to interpret. Structured output is handing them a form with labelled boxes: vendor, date, amount, currency. A schema is the form's design; validation is the clerk who checks every box is filled in properly before the form goes into the filing system.

Examples

Parse, validate and report errors (offline)

// A minimal schema validator, enough to show the idea
const schema = {
  category: { type: "string", enum: ["shipping", "billing", "account", "other"] },
  urgent: { type: "boolean" },
  summary: { type: "string", maxLength: 80 },
};

function validate(obj, schema) {
  const errors = [];
  for (const [field, rule] of Object.entries(schema)) {
    const v = obj[field];
    if (v === undefined) { errors.push(field + ": missing"); continue; }
    if (typeof v !== rule.type) errors.push(field + ": expected " + rule.type + ", got " + typeof v);
    if (rule.enum && !rule.enum.includes(v)) errors.push(field + ": must be one of " + rule.enum.join("|"));
    if (rule.maxLength && typeof v === "string" && v.length > rule.maxLength) errors.push(field + ": too long");
  }
  for (const key of Object.keys(obj)) if (!(key in schema)) errors.push(key + ": unexpected field");
  return errors;
}

// Strip a markdown code fence if the model wrapped its JSON in one
const FENCE = String.fromCharCode(96).repeat(3);
function extractJson(text) {
  let t = text.trim();
  if (t.startsWith(FENCE)) t = t.replace(/^\S*\s*/, "").replace(new RegExp(FENCE + "\\s*
quot;), ""); return JSON.parse(t); } const modelOutputs = [ '{"category": "billing", "urgent": true, "summary": "Charged twice"}', FENCE + 'json {"category": "billing", "urgent": "yes", "summary": "Charged twice"} ' + FENCE, '{"category": "payments", "urgent": false}', 'Sure! Here is the JSON you asked for: {"category": "other"}', ]; for (const out of modelOutputs) { try { const obj = extractJson(out); const errors = validate(obj, schema); console.log(errors.length ? "INVALID: " + errors.join("; ") : "VALID: " + JSON.stringify(obj)); } catch (e) { console.log("NOT JSON:", e.message.slice(0, 60)); } }

Four realistic outputs from 'just asking for JSON': one good, one wrapped in a code fence with a wrong type, one with an invalid enum value and a missing field, and one with chatty text around it. This is why you validate - or better, use the API's structured-output support.

Validate-and-retry loop with a fake model

// The fake model gets it wrong first, then fixes it when shown the error.
let attempt = 0;
function fakeModel(prompt) {
  attempt++;
  if (attempt === 1) return '{"category": "Billing", "urgent": "true"}';
  return '{"category": "billing", "urgent": true}';
}

function validate(obj) {
  const errors = [];
  if (!["shipping", "billing", "account", "other"].includes(obj.category)) errors.push("category must be one of shipping|billing|account|other");
  if (typeof obj.urgent !== "boolean") errors.push("urgent must be a boolean true/false, not a string");
  return errors;
}

function getStructured(task, maxAttempts = 3) {
  let prompt = task;
  for (let i = 1; i <= maxAttempts; i++) {
    const raw = fakeModel(prompt);
    let errors;
    try { errors = validate(JSON.parse(raw)); } catch (e) { errors = ["invalid JSON: " + e.message]; }
    console.log("attempt", i, raw, errors.length ? "-> " + errors.join("; ") : "-> ok");
    if (errors.length === 0) return JSON.parse(raw);
    prompt = task + " Your previous answer was invalid: " + errors.join("; ") + ". Return corrected JSON only.";
  }
  throw new Error("no valid output after " + maxAttempts + " attempts");
}

console.log("result:", JSON.stringify(getStructured("Classify: 'I was charged twice!'")));

Feeding the specific validation error back is much more effective than simply asking again. Always cap the attempts and have a fallback (default value, human review) when the cap is reached.

Structured outputs with Pydantic and messages.parse (Python)

from typing import Literal, Optional
import anthropic
from pydantic import BaseModel, Field

class Invoice(BaseModel):
    vendor: str
    invoice_number: str
    total: float = Field(description="Total amount including tax")
    currency: Literal["USD", "EUR", "GBP", "OTHER"]
    due_date: Optional[str] = Field(None, description="ISO date YYYY-MM-DD, or null if not stated")

client = anthropic.Anthropic()

text = """INVOICE #A-1042 from Northwind Supplies
Total due: 1,250.00 EUR (incl. VAT). Payment within 14 days."""

resp = client.messages.parse(
    model="claude-opus-5-5",
    max_tokens=16000,
    messages=[{"role": "user", "content": f"Extract the invoice fields from:\n<invoice>{text}</invoice>"}],
    output_format=Invoice,
)

inv = resp.parsed_output          # an Invoice instance
print(inv.vendor, inv.invoice_number, inv.total, inv.currency, inv.due_date)

# Structure is guaranteed; business rules are still your job:
if inv.total <= 0:
    raise ValueError("total must be positive")

The Pydantic model is converted to a JSON Schema and sent with the request; the output is constrained to it and parsed back into a typed Python object. Note due_date is optional: the invoice gives '14 days', not a date, so None is the honest answer rather than an invented date.

Validate-and-retry with Pydantic (works with any model)

import anthropic
from pydantic import BaseModel, ValidationError

class Ticket(BaseModel):
    category: str
    urgent: bool

client = anthropic.Anthropic()
messages = [{"role": "user", "content":
    'Classify this ticket. Reply with only JSON like {"category": "...", "urgent": true}.\n'
    "<ticket>Site is down for everyone!</ticket>"}]

ticket = None
for attempt in range(3):
    resp = client.messages.create(model="claude-opus-5-5", max_tokens=16000, messages=messages)
    raw = next(b.text for b in resp.content if b.type == "text")
    try:
        ticket = Ticket.model_validate_json(raw)
        break
    except ValidationError as e:
        messages.append({"role": "assistant", "content": raw})
        messages.append({"role": "user", "content": f"That was invalid: {e}. Reply with corrected JSON only."})

print(ticket)

This is the portable pattern: parse, validate, and send the validation error back. Prefer messages.parse when available because it removes most failures up front.

How it works

Prompt-only JSON relies on the model following instructions; it is usually right but has no guarantee, so parsing can fail.

Constrained decoding (what API structured-output features do under the hood) restricts which tokens the model may choose at each step so the text being generated always stays a valid prefix of something matching the schema. If the schema says the next thing must be true or false, tokens that would start anything else are excluded. The result always parses and matches the structure.

The SDK's parse helper converts your Pydantic class to JSON Schema, sends it with the request, and validates the response back into an instance. Some JSON Schema features may not be supported by the constraint engine (for example certain numeric or string-length constraints); keep schemas simple and enforce extra rules in your own validation.

Validation layers: (1) syntax - is it JSON? (2) structure - required fields and types (schema); (3) semantics - business rules and cross-field checks (your code); (4) truth - does the value match the source (spot checks, evaluation).

   input text
       |
       v
 model + schema -----> constrained output
       |                (always parses)
       v
 parse JSON ------ fail --> retry with error
       |
 schema validate - fail --> retry with error
       |
 business rules -- fail --> flag / human
       |
       v
  typed object -> database / API / UI

Why does it exist?

Integrating LLMs into software requires outputs that code can rely on. Free text breaks parsers in production at the worst moments. Schemas, validation and constrained decoding turn an LLM from a text generator into a dependable component: an extractor, classifier or router with a typed interface.

When to use it

Use structured output whenever code consumes the result: classification, extraction from documents, routing decisions, generating records, filling forms, producing tool arguments, scoring in evaluations. Use enums for anything that maps to fixed options in your system.

When not to use it

Do not force JSON for content meant for humans (an email reply, an explanation) - it adds escaping noise and can lower writing quality. For long prose plus a little metadata, return the prose as a string field or use tags. Do not build huge, deeply nested schemas for a single call; split the task.

Common mistakes

  • Calling JSON.parse/json.loads on raw model text with no error handling.

  • Not stripping markdown code fences or surrounding prose when using prompt-only JSON.

  • Making every field required, forcing the model to invent values that are not in the source.

  • Assuming schema-valid means correct; values can still be wrong.

  • Using free-text fields where an enum of allowed values belongs.

  • Retrying without telling the model what was wrong.

  • Retrying forever instead of capping attempts and falling back.

Practice exercises

  1. Easy:

    Write a JSON Schema (or Pydantic model) for a contact card: name (required), email (required), phone (optional), company (optional).

  2. Easy:

    Add an items array field to the offline validator demo's schema and support validating that it is an array of strings.

  3. Medium:

    Extend extractJson to also handle prose before and after the JSON by locating the first '{' and the matching last '}'. Test on the four sample outputs.

  4. Medium:

    Use messages.parse to extract a list of action items (owner, task, due date or null) from a meeting transcript. Add business-rule validation that every owner is in a known list of attendees.

  5. Hard:

    Build a batch extractor: given 20 short job postings, extract title, company, salary range (nullable numbers), remote (enum: yes/no/hybrid/unknown). Validate, count failures by field, and write a CSV of the results.

Interview questions

How do you get reliable JSON from an LLM?

Use the API's structured-output feature with a schema (e.g. messages.parse with a Pydantic model), which constrains generation to valid output. Otherwise, include the schema and examples in the prompt, parse and validate, and retry with the error message. Always validate business rules in code.

What is constrained decoding?

Restricting the tokens the model may choose at each step so the output always remains valid according to a grammar or schema. It guarantees the structure parses, but not that the values are correct.

Why make some schema fields nullable?

If a field is required but the information is absent from the input, the model is pushed to invent a value. Allowing null lets it report 'not present' honestly.

Why put a reasoning field before the answer fields?

Output is generated in order, so a reasoning field first lets the model work through the problem before committing to the answer fields, which can improve accuracy.

What validation do you still need with guaranteed schema conformance?

Semantic checks: value ranges, cross-field consistency, references to real entities, and spot-checking that extracted values actually appear in the source.