Generating Into a Schema: When Prose Is the Wrong AI Output Format
Free prose is the easy AI output. When it needs to be rendered, stored, revised, or reproduced consistently, a schema is what makes it a product.
The demo generates beautiful prose. The production system cannot use it.
Prose reads well and looks compelling in a screenshot. It also cannot be rendered into a specific UI component, compared across two runs, stored in a way that lets individual fields be updated, or fed into a downstream process that expects predictable structure. When AI output needs to do any of those things, prose is the wrong format.
Why a schema is the real specification
Asking an AI to “generate a relocation plan” produces something. Asking it to “fill this schema with a personalised plan” produces something you can build on.
The schema is what makes the output a deliverable. It defines exactly what the generation task is: which fields need values, which are optional, what type each field holds, and what null means versus absent. Before writing a single prompt, writing the schema forces a decision about what the output actually is. If you cannot define the schema, you do not yet know what the AI is supposed to produce.
This is the same point behind scoping an AI MVP: “summarise this document” is a task with no definition of done. “Extract these fields into this schema” has one. The difference between a demo and a product is often just that one.
What structured generation looks like in practice
Most model providers support structured output modes that return JSON conforming to a schema you supply. The generation prompt becomes: given this input, fill these fields with the right values.
The result is checkable. Does the output conform to the schema? This is a binary question. Whether prose captures the right meaning is not. Fields can be null when information is absent, which is honest. Prose hides absence behind vague sentences or omits it silently.
Arrays of items in a schema are indexable, searchable, and renderable as lists. An array of relocation steps, each with a label, a description, and a set of actions, can be turned into a PDF section, a UI list, or a database row. The equivalent in free prose is a paragraph that a human can read but code cannot process.
Designing the schema to match the output format
The schema should reflect how the output will be used, not just what it contains.
When the deliverable is a document, the schema should mirror the document’s sections. A personalised plan has sections, and each section has items, and each item has a heading, a body, and optional sub-items. The schema mirrors that structure exactly, so rendering the document is a matter of iterating the schema rather than parsing prose.
When the output feeds a UI, the schema should mirror the UI components. A card has a title, a subtitle, a body, and a status field. The AI fills those fields. The UI renders them. There is no parsing step in the middle.
Flat schemas are simpler to validate and render. Deeply nested schemas are sometimes necessary but should be questioned: every additional nesting level is a place where optional fields and null handling get complicated.
Avoid fields defined as “anything else” or “additional context.” They are just prose with a label, and they inherit all of prose’s problems. If a field’s value cannot be described by a type, it probably belongs in the schema as its own field.
The artefact changes the design
When you know the output is a PDF, you can design the schema to match how a PDF is structured: sections with headings, paragraphs, lists, and metadata at the top level. The generation produces values. The renderer consumes them. The separation is clean.
This was the central design decision behind the relocation planning tool we built for Olim Paveway. The product’s real output is a PDF plan, and the schema was designed around that from the start: each stage of the relocation process is a structured item with a heading, a description, a list of actions, and a timeline. Generating into that schema is what makes the plan renderable as a document, revisable without regenerating the whole thing, and consistent between runs rather than stylistically different each time.
The same discipline appears in RFP response automation, where every requirement, evaluation criterion, and submission deadline is extracted into a defined schema rather than a summary. The schema is what makes the extraction checkable against something.
When free prose is the right choice
Short explanatory text that sits alongside structured data is often better as prose. A one-sentence summary accompanying a structured result, an explanation of why a field has a particular value, a human-readable description that will be read directly and never processed further.
The test: will code ever need to do something specific with this output? If yes, give it a schema. If the only thing that will happen is a human reads it, prose is fine.
The practical version: if the output goes into a component, a document template, a database column, or a downstream API call, it needs a schema. If it goes into a text box that a human reads and discards, prose is fine.
We built the Olim Paveway plan generator this way. See the case study, or read about our approach to AI agent work.