Most LLM features in real products do not end with text on a screen. They end with data: a ticket category, fields extracted from an invoice, a list of action items, arguments for an API call. For that, "please respond in JSON" is not enough. Even a 1% parse failure rate means hundreds of broken requests a day at modest scale. This guide covers the mechanisms that make structured output reliable, how to design schemas that models fill correctly, and how to validate and recover when something still goes wrong.

Three ways to get structured output

Approach How it works Reliability
Prompting only "Respond with JSON matching this example" Good with modern models, but not guaranteed
Tool / function calling You define a tool with a JSON schema; the model returns arguments High; widely supported
Structured output / constrained decoding The API constrains generation so output must match the schema Highest — syntactically guaranteed

Major providers now support schema-constrained output natively (OpenAI structured outputs; Anthropic structured outputs). For self-hosted models, libraries such as Outlines and inference servers like vLLM implement constrained decoding by masking tokens that would violate the grammar (Willard & Louf, 2023).

Use constrained output whenever you can. It eliminates an entire class of failures — unparseable JSON, missing required fields, wrong types — and lets you focus on the hard part: whether the values are correct.

Syntax is guaranteed; semantics are not

Constrained decoding guarantees the output parses and matches the schema. It does not guarantee that the invoice total is the right number, that the category is correct, or that the date is not invented. A model forced to fill a required field will fill it — even when the information is not in the input. This is the most important thing to understand about structured output.

Design for it:

  • Make missing information representable. Use nullable fields or explicit "unknown" values instead of forcing a guess.
  • Validate business rules in code: totals equal the sum of lines, dates are in a plausible range, IDs exist in your database.
  • Evaluate value accuracy on a labelled dataset, as you would any other LLM feature; see LLM evals.

Schema design that models fill correctly

The schema is part of the prompt. Models read field names, descriptions and enum values. Good schemas follow a few rules:

  1. Descriptive field names. invoice_total_with_vat beats amt2.
  2. Descriptions on every non-obvious field, including units and formats: "ISO 8601 date", "amount in UAH, decimal, no currency symbol".
  3. Enums for closed sets of values, with an "other" option when the world is messier than your list.
  4. Flat over deeply nested. Two levels of nesting are fine; five are error-prone.
  5. Order fields to support reasoning. Put a reasoning or evidence field before the decision field if you need the model to justify a classification; generation is left to right.
  6. Avoid huge enums. For hundreds of categories, retrieve candidates first and give the model a short list.
from enum import Enum
from pydantic import BaseModel, Field

class Category(str, Enum):
    billing = "billing"
    technical = "technical"
    account = "account"
    other = "other"

class TicketTriage(BaseModel):
    evidence: str = Field(description="Quote the sentence(s) that justify the category.")
    category: Category
    urgency: int = Field(ge=1, le=3, description="1 = low, 3 = service down for customer")
    customer_order_id: str | None = Field(
        default=None,
        description="Order ID like ORD-123456 if explicitly mentioned, otherwise null.",
    )

Validation and recovery

Even with constrained decoding, validate every response in your application code. Pydantic in Python and Zod in TypeScript let you define the schema once and use it for both the API request and validation:

import { z } from "zod"

const LineItem = z.object({
  description: z.string(),
  quantity: z.number().positive(),
  unitPrice: z.number().nonnegative(),
})

export const Invoice = z.object({
  supplierName: z.string(),
  supplierTaxId: z.string().regex(/^\d{8,10}$/).nullable(),
  issueDate: z.string().date(),
  currency: z.enum(["UAH", "EUR", "USD"]),
  lines: z.array(LineItem).min(1),
  total: z.number().nonnegative(),
}).refine(
  (inv) => Math.abs(inv.lines.reduce((s, l) => s + l.quantity * l.unitPrice, 0) - inv.total) < 0.01,
  { message: "total does not match sum of lines" },
)

When validation fails, you have options:

  1. Retry with the error. Send the model its output plus the validation message: "total does not match sum of lines — re-check the line items". One retry fixes most issues; more than two rarely helps.
  2. Fall back to a stronger model for the failed item.
  3. Route to a human with the partial result pre-filled. For document processing, this is often the right default; see LLM document processing.

Libraries such as Instructor wrap this validate-and-retry loop for many providers. The Vercel AI SDK offers generateObject with Zod schemas for TypeScript; we use it in building AI chat in Next.js.

Does forcing JSON hurt quality?

Some research found that strict format restrictions can degrade reasoning performance on some tasks compared with free-form answers (Tam et al., 2024). In practice the effect depends on the model and the task, and modern structured-output implementations have narrowed the gap. Two mitigations work well:

  • Reason first, then structure. Let the model think (extended thinking, or a free-text reasoning field placed before the answer fields), then emit the structured result.
  • Two-step pipelines for hard tasks: one call produces a free-form analysis, a second, cheaper call extracts the structured fields from it.

Measure on your own eval set rather than assuming either way.

Tool calling as structured output

Tool calling and structured output are two faces of the same mechanism: the model emits arguments that match a schema. Use tool calling when the model should decide whether and which action to take; use structured output when you always want one specific shape. In agents, every tool's input schema deserves the same care as the schemas above — see designing tools for LLM agents.

Streaming structured output

For user-facing features, you may want to show partial results as they are generated: a list of action items appearing one by one, a form filling progressively. Many SDKs support streaming partial objects; the client receives incomplete JSON that is "repaired" into a valid partial object. Validate only the final object, and design UI that tolerates fields appearing and changing.

Performance and cost tips

  • Schemas cost tokens. Large schemas with long descriptions add input tokens to every call; they are stable, so prompt caching makes them cheap — see LLM cost optimization.
  • First-call latency. Some providers compile the schema into a grammar on first use; expect a one-time delay for new schemas.
  • Batch extraction. For offline jobs (thousands of documents), batch APIs cut costs substantially.
  • Smaller models are often enough for well-defined extraction with good schemas. Evaluate before defaulting to the largest model.

Measuring extraction quality

Structured output makes evaluation easier than free text, because you can compare fields directly. Build a labelled set of 100–300 real inputs and measure per field:

Metric What it tells you
Field accuracy Share of exact (or normalised) matches per field
Null precision When the model returns null, was the value really absent?
Null recall When the value was absent, did the model return null instead of inventing one?
Document-level accuracy Share of inputs where all critical fields are correct
Validation failure rate How often business rules reject the output

Null recall is the metric most teams forget and the one that reveals hallucinated values. Document-level accuracy is what matters for automation decisions: if 92% of invoices are fully correct, you can auto-process those that pass validation and send the rest to review.

Common mistakes

  • Required fields for optional information, causing invented values.
  • Trusting enums blindly. A correct-looking category can still be wrong; measure accuracy per class.
  • No business validation. Schema-valid does not mean correct.
  • Parsing JSON out of markdown with regex in 2026. Use the native structured output features.
  • One giant schema for everything. Split extraction into focused calls when the schema grows beyond what a human could fill in one pass.

FAQ

Is JSON mode the same as structured outputs? No. Older "JSON mode" guarantees valid JSON but not a specific schema. Schema-constrained structured outputs guarantee both.

What about XML or YAML? JSON with a schema is the best-supported option. XML tags remain useful for free-text sections inside prompts and responses.

How do we handle very long lists? Paginate or chunk the input, extract per chunk, and merge in code. Long outputs increase the chance of truncation and omission.

Can we use structured output with self-hosted models? Yes. vLLM and similar servers support guided decoding with JSON schemas; see self-hosting LLMs.

Sources

  1. OpenAI. Structured outputs and function calling.
  2. Anthropic. Structured outputs and tool use.
  3. Willard, B., Louf, R. (2023). Efficient Guided Generation for Large Language Models.
  4. Tam et al. (2024). Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models.
  5. JSON Schema, Pydantic, Zod, Instructor.