By the end of this post you will be able to look at a prompt that keeps producing malformed JSON and name the specific structural reason it fails, then apply a targeted fix instead of regenerating until it works. The reasons are not mysterious. They fall into a small number of categories, and each one has a countermeasure you can write into the prompt itself.
The post is organized as a beginner versus advanced comparison. The first half covers the mistakes that account for most broken output in single-turn, low-stakes use. The second half covers the failure modes that appear once you move to programmatic pipelines, schema constraints, and retries — the ones that survive a naive prompt fix and need a different layer of control.
Part 1: Beginner Mistakes (Single Prompt, Manual Use)
If you are pasting a prompt into a chat interface and copying JSON out by hand, these are the failure modes you will hit most often. Each fix is a change you make inside the prompt text.
Mistake 1: Asking for JSON without saying what shape it should be
The most common cause of unusable output is a request like “return the results as JSON” with no schema. The model has to guess field names, nesting depth, and whether you want an array or an object. It will pick something plausible, and plausible is not the same as compatible with your parser.
A weak prompt:
Extract the key information from this support ticket as JSON.
Ticket: "Hi, I was charged twice for order #88421 on May 3. My card ends in 4417.
Please refund the duplicate."
The model might return {"order": "88421", "issue": "double charge"} or {"order_id": 88421, "problem": {...}} or a plain string. Any of these is a reasonable interpretation, and only one of them matches your downstream code.
The fix is to print the schema, key by key, inside the prompt. Do not describe it in prose. Show it.
Extract the key information from this support ticket. Return only a JSON object
matching this exact schema:
{
"order_id": string,
"charge_date": string (ISO 8601, YYYY-MM-DD),
"card_last_four": string,
"issue_type": one of ["double_charge", "late_delivery", "wrong_item", "other"],
"requested_action": string
}
Ticket: "Hi, I was charged twice for order #88421 on May 3. My card ends in
4417. Please refund the duplicate."
A shown schema does three things a prose description does not. It fixes the key names, it communicates the expected types, and it makes the enum values concrete so the model has a closed set to choose from rather than inventing a label. The word “exact” matters — it signals that deviation is not acceptable.
Mistake 2: Leaving the JSON inside prose
A model that answers in sentences will sometimes wrap JSON in a friendly sentence: “Sure, here is the extracted data: { … } Let me know if you need anything else.” Your parser sees the leading text, calls JSON.parse, and throws. The JSON itself was correct; the surrounding prose broke it.
You cannot rely on the model to always omit the preamble. Instead, state the output constraint explicitly and give it a reason not to narrate. Two effective phrasings:
- “Return only the JSON object. Do not include any text before or after it.”
- “Output must be machine-parseable with no additional commentary.”
For stricter cases, add a negative example. Showing the model what not to do is often more reliable than a single instruction, because it demonstrates the boundary rather than describing it:
Return only raw JSON. Do not wrap it in markdown code fences, do not add a
sentence before or after it.
Correct: {"status": "ok", "count": 3}
Incorrect: Here is the JSON: {"status": "ok", "count": 3}
Incorrect: ```json
{"status": "ok", "count": 3}
### Mistake 3: Forgetting that markdown fences are a common default
Many chat models default to wrapping code-like output in triple backticks with a `json` language tag. That is correct markdown and incorrect for a raw parser. If your pipeline reads the model's text output directly, a fenced block will fail at the same place a prose wrapper does.
Name the constraint directly: "Do not wrap the JSON in markdown code fences." If you cannot prevent the fence, plan to strip it. A small extraction step that finds the first `{` and last `}` and slices between them handles both the fence and the prose-wrapper case at once:
```python
import json
def extract_json(text: str) -> dict:
"""Pull the first JSON object out of a model response."""
start = text.find("{")
end = text.rfind("}")
if start == -1 or end == -1 or end < start:
raise ValueError("No JSON object found in response")
return json.loads(text[start:end + 1])
This function is a defensive layer, not a substitute for a good prompt. It also has a real limitation: if the model produces two JSON objects, or an array instead of an object, the slice returns the wrong thing or nothing. Treat it as a recovery path for occasional messy responses, not the primary contract.
Mistake 4: Not specifying types, so numbers arrive as strings
Without type information, a model may return "order_id": "88421" when your code expects an integer, or "charge_date": "May 3" when you need 2026-05-03. Both parse as valid JSON, so validation passes and the bug shows up later in a comparison or date parse.
State the type in the schema and, for dates and numbers, state the format. A schema line like "charge_date": string in ISO 8601 format (YYYY-MM-DD) leaves nothing to interpret. For booleans and counts, add the constraint directly: "is_duplicate": boolean, "refund_amount_cents": integer.
Mistake 5: Using vague enum values or open-ended categories
If a field is meant to hold one of a fixed set of values, list that set. An open field like "category": string invites the model to produce "billing", "Billing Issue", "billing-related", and "payment problem" for what you intended as one category. Your downstream if category == "billing" check then fails on three of those four.
Enumerate the values and use the word “one of”:
"category": one of ["billing", "shipping", "product_quality", "account_access"]
This is the single highest-leverage change for classification tasks. It converts an open generation problem into a constrained selection problem, and the model is typically far more consistent at selection than at invention.
Beginner Fix Summary
| Mistake | Symptom | Prompt-level Fix |
|---|---|---|
| No schema | Wrong key names, arbitrary nesting | Print the exact JSON shape in the prompt |
| Prose wrapper | Parser fails on leading text | “Return only JSON, no text before or after” |
| Markdown fence | Backticks break parsing | “Do not wrap in code fences”; strip as backup |
| Missing types | Numbers as strings, wrong date format | State type and format per field |
| Open enum | Inconsistent category values | List allowed values with “one of” |
Part 2: Advanced Mistakes (Programmatic Pipelines and Structured Output)
The beginner fixes clean up most single-turn problems. They do not solve the ones that appear when JSON output becomes a contract your code depends on, running thousands of times with no human in the loop. Those require controls outside the prompt text.
The intermediate technique that changes the game: schema-constrained decoding
Before listing the advanced mistakes, it helps to understand the mechanism that removes an entire class of them. Several major providers now offer a mode where you pass a JSON schema (typically in JSON Schema draft format) alongside the prompt, and the model is constrained to emit only tokens that keep the output valid against that schema. OpenAI exposes this as response_format with type: "json_schema" and a strict: true flag. Anthropic exposes a comparable capability through tool definitions where the tool input schema is enforced. Google’s Gemini API offers a controlled generation mode with a responseSchema field.
When you use these modes, the model cannot return invalid JSON — the sampler is not allowed to produce the closing brace early or emit a bare unquoted key. That does not mean your data is correct. It means the syntax is guaranteed and the semantics are still yours to check. A schema can force a field to be a string, but it cannot force that string to be a real email address.
If the option is available to you, use it. It removes the fence problem, the prose problem, and the closed-brace problem in one step. The rest of this section covers what you still have to handle after that.
Mistake 6: Schema drift between prompt and validator
When the schema lives only in the prompt, it can silently disagree with the schema your validator enforces. Someone updates a field name in the prompt and forgets the Pydantic model, or adds a field to the validator that the prompt never mentions. The failure is intermittent and depends on which fields the model fills.
The fix is to make the schema a single source of truth and generate the prompt text from it. Define the schema once — in JSON Schema, Pydantic, or a typed struct — then serialize it into the prompt. This is the one change that eliminates drift by construction.
from pydantic import BaseModel, Field
from typing import Literal
class Ticket(BaseModel):
order_id: str
charge_date: str = Field(pattern=r"^\d{4}-\d{2}-\d{2}$")
card_last_four: str
issue_type: Literal["double_charge", "late_delivery", "wrong_item", "other"]
requested_action: str
schema_json = Ticket.model_json_schema()
prompt = f"""Extract the ticket fields. Return only JSON matching this schema:
{schema_json}
"""
Now the schema the model sees and the schema you validate against come from the same object. Change the model, and both update.
Mistake 7: No validation step, so invalid data flows downstream
Even with a perfect prompt, a model will occasionally return a value that is the right type but the wrong content — a date in the wrong format, a category outside the enum, a missing required field. Without a validation step, that data reaches your database or your UI and causes a failure far from where it was created.
Validate every response before using it. Pydantic raises a clear error, and you can decide whether to retry, log, or fall back:
from pydantic import ValidationError
def parse_ticket(raw: str) -> Ticket | None:
try:
data = extract_json(raw)
return Ticket.model_validate(data)
except (ValueError, ValidationError) as exc:
# Log the raw response and the error for later inspection
print(f"Validation failed: {exc}\nRaw: {raw}")
return None
The value here is not just catching errors early. It is that the validation error tells you what the model got wrong, which is the signal you feed back into a retry.
Mistake 8: Retrying with the same prompt
When validation fails, the instinct is to call the model again with the same prompt and hope for a different result. That works sometimes, but it wastes the information the validator already produced. The error message is a precise description of what was wrong, and you can put it back into the conversation.
A targeted retry appends the validation error and asks for a correction:
def parse_with_retry(raw: str, prompt: str, client, max_attempts: int = 3):
for attempt in range(max_attempts):
result = parse_ticket(raw)
if result is not None:
return result
# Send the error back and request only a corrected object
correction = (
f"The previous output failed validation with this error:\n"
f"{last_error}\n"
f"Return a corrected JSON object only. Do not explain the error."
)
response = client.complete(prompt=prompt, messages=[correction])
raw = response.text
raise RuntimeError("Exceeded retry budget without valid output")
Two details matter. First, the retry asks for the corrected object and explicitly forbids explanation, which keeps the response parseable. Second, the retry budget is bounded. An unbounded retry loop on a prompt that is fundamentally underspecified will burn tokens and time without converging.
Mistake 9: Assuming structured output means correct output
Schema-constrained decoding guarantees syntax, not truth. A model can return a perfectly valid object with "charge_date": "1901-01-01" because the schema only required a date-shaped string. The syntax passed; the content is wrong.
This is where the advanced techniques end and domain validation begins. For a field like a date or an amount, validate the value, not just the shape. For a field like a category, confirm membership in the enum. For cross-field consistency, check the relationships your schema cannot express — a refund amount that exceeds the original charge, a start date after an end date. Schema constraints handle the container. Business rules handle what goes inside it.
Mistake 10: Over-constraining and forcing the model to hallucinate
There is a real cost to pushing too hard on structure. If you demand a field that the input does not contain, the model has three options: leave it null, omit it, or invent a plausible value. Many prompts discourage nulls, and omission violates a required schema, so the model fills the gap. It fabricates.
A ticket that mentions no order ID but requires one will produce a confabulated number that looks correct and is not. The fix is to make optionality explicit. Mark fields that may be absent, and in the schema allow null. Instruct the model directly: “If a field is not present in the input, use null. Do not guess.” Then make your downstream code handle null rather than assuming presence.
Over-constraining also shows up in enum design. A closed set of five categories that does not include the input’s real category forces the model to shoehorn it into the closest wrong option. Always include an “other” or “unclassified” value.
Advanced Fix Summary
| Mistake | Failure Mode | Fix |
|---|---|---|
| Schema drift | Prompt and validator disagree | Generate prompt from one schema object |
| No validation | Bad data reaches downstream | Validate every response with Pydantic |
| Blind retry | Same error repeats; wasted calls | Feed the validation error back, bound attempts |
| Trusting valid syntax | Structurally valid, factually wrong | Add domain and cross-field checks |
| Over-constraining | Model hallucinates missing fields | Allow nulls; include an “other” category |
Beginner vs Advanced: When to Use Which
Prompt-level fixes are sufficient when a human reviews the output, the volume is low, and a failed parse costs a copy-paste. In those settings, a clear schema in the prompt and an instruction to omit prose will handle almost everything.
Move to schema-constrained decoding and a validation layer when the output feeds code with no human in between, when the same prompt runs at volume, or when a wrong value has a cost beyond annoyance. The trade-off is real: constrained decoding limits flexibility, adds a schema to maintain, and in some providers can reduce the model’s ability to reason through the response before emitting it. For a task where you want the model to think first and commit to JSON last, forcing structured output from the first token can degrade quality. In those cases a two-step approach works better — reason in free text, then emit JSON in a second constrained call.
Do not add validation and retries to a prompt that is simply missing a schema. The simplest fix is usually the prompt text itself, and the advanced machinery is there to handle the residual failures, not to compensate for an underspecified request.
A Starting Point
If you take one thing from this post, make it the schema. Print the exact JSON shape you expect inside the prompt, generate it from the same object you validate against, and name the types, formats, and enum values explicitly. That single discipline removes the majority of the beginner mistakes. Add validation and targeted retries when you go programmatic, and reserve schema-constrained decoding for the cases where structural guarantees are worth the reduced flexibility. The task is to make the gap between what you asked for and what your parser expects as small as possible — and then to check, rather than assume, that the model stayed inside it.
🔗 Recommended Reading
- How to Evaluate LLM Outputs: A Beginner's Step-by-Step Tutorial to Testing and Scoring AI Responses
- LLM Guardrails for Beginners: A Step-by-Step Tutorial to Filtering Unsafe Outputs
- Building Your First AI Agent: A Beginner Tutorial with Python and the OpenAI API
- Function Calling in LLM APIs: A Beginner Tutorial for Connecting AI to Real Tools
- RAG Chunking Strategies for Beginners: A Step-by-Step Tutorial