By the end of this post you will be able to look at a prompt that keeps producing malformed JSON and name the specific structural reason it fails, then apply a targeted fix instead of regenerating until it works. The reasons are not mysterious. They fall into a small number of categories, and each one has a countermeasure you can write into the prompt itself.

The post is organized as a beginner versus advanced comparison. The first half covers the mistakes that account for most broken output in single-turn, low-stakes use. The second half covers the failure modes that appear once you move to programmatic pipelines, schema constraints, and retries — the ones that survive a naive prompt fix and need a different layer of control.

Part 1: Beginner Mistakes (Single Prompt, Manual Use)

If you are pasting a prompt into a chat interface and copying JSON out by hand, these are the failure modes you will hit most often. Each fix is a change you make inside the prompt text.

Mistake 1: Asking for JSON without saying what shape it should be

The most common cause of unusable output is a request like “return the results as JSON” with no schema. The model has to guess field names, nesting depth, and whether you want an array or an object. It will pick something plausible, and plausible is not the same as compatible with your parser.

A weak prompt:

Extract the key information from this support ticket as JSON.

Ticket: "Hi, I was charged twice for order #88421 on May 3. My card ends in 4417.
Please refund the duplicate."

The model might return {"order": "88421", "issue": "double charge"} or {"order_id": 88421, "problem": {...}} or a plain string. Any of these is a reasonable interpretation, and only one of them matches your downstream code.

The fix is to print the schema, key by key, inside the prompt. Do not describe it in prose. Show it.

Extract the key information from this support ticket. Return only a JSON object
matching this exact schema:

{
  "order_id": string,
  "charge_date": string (ISO 8601, YYYY-MM-DD),
  "card_last_four": string,
  "issue_type": one of ["double_charge", "late_delivery", "wrong_item", "other"],
  "requested_action": string
}

Ticket: "Hi, I was charged twice for order #88421 on May 3. My card ends in
4417. Please refund the duplicate."

A shown schema does three things a prose description does not. It fixes the key names, it communicates the expected types, and it makes the enum values concrete so the model has a closed set to choose from rather than inventing a label. The word “exact” matters — it signals that deviation is not acceptable.

Mistake 2: Leaving the JSON inside prose

A model that answers in sentences will sometimes wrap JSON in a friendly sentence: “Sure, here is the extracted data: { … } Let me know if you need anything else.” Your parser sees the leading text, calls JSON.parse, and throws. The JSON itself was correct; the surrounding prose broke it.

You cannot rely on the model to always omit the preamble. Instead, state the output constraint explicitly and give it a reason not to narrate. Two effective phrasings:

  • “Return only the JSON object. Do not include any text before or after it.”
  • “Output must be machine-parseable with no additional commentary.”

For stricter cases, add a negative example. Showing the model what not to do is often more reliable than a single instruction, because it demonstrates the boundary rather than describing it:

Return only raw JSON. Do not wrap it in markdown code fences, do not add a
sentence before or after it.

Correct: {"status": "ok", "count": 3}
Incorrect: Here is the JSON: {"status": "ok", "count": 3}
Incorrect: ```json
{"status": "ok", "count": 3}

### Mistake 3: Forgetting that markdown fences are a common default

Many chat models default to wrapping code-like output in triple backticks with a `json` language tag. That is correct markdown and incorrect for a raw parser. If your pipeline reads the model's text output directly, a fenced block will fail at the same place a prose wrapper does.

Name the constraint directly: "Do not wrap the JSON in markdown code fences." If you cannot prevent the fence, plan to strip it. A small extraction step that finds the first `{` and last `}` and slices between them handles both the fence and the prose-wrapper case at once:

```python
import json

def extract_json(text: str) -> dict:
    """Pull the first JSON object out of a model response."""
    start = text.find("{")
    end = text.rfind("}")
    if start == -1 or end == -1 or end < start:
        raise ValueError("No JSON object found in response")
    return json.loads(text[start:end + 1])

This function is a defensive layer, not a substitute for a good prompt. It also has a real limitation: if the model produces two JSON objects, or an array instead of an object, the slice returns the wrong thing or nothing. Treat it as a recovery path for occasional messy responses, not the primary contract.

Mistake 4: Not specifying types, so numbers arrive as strings

Without type information, a model may return "order_id": "88421" when your code expects an integer, or "charge_date": "May 3" when you need 2026-05-03. Both parse as valid JSON, so validation passes and the bug shows up later in a comparison or date parse.

State the type in the schema and, for dates and numbers, state the format. A schema line like "charge_date": string in ISO 8601 format (YYYY-MM-DD) leaves nothing to interpret. For booleans and counts, add the constraint directly: "is_duplicate": boolean, "refund_amount_cents": integer.

Mistake 5: Using vague enum values or open-ended categories

If a field is meant to hold one of a fixed set of values, list that set. An open field like "category": string invites the model to produce "billing", "Billing Issue", "billing-related", and "payment problem" for what you intended as one category. Your downstream if category == "billing" check then fails on three of those four.

Enumerate the values and use the word “one of”:

"category": one of ["billing", "shipping", "product_quality", "account_access"]

This is the single highest-leverage change for classification tasks. It converts an open generation problem into a constrained selection problem, and the model is typically far more consistent at selection than at invention.

Beginner Fix Summary

MistakeSymptomPrompt-level Fix
No schemaWrong key names, arbitrary nestingPrint the exact JSON shape in the prompt
Prose wrapperParser fails on leading text“Return only JSON, no text before or after”
Markdown fenceBackticks break parsing“Do not wrap in code fences”; strip as backup
Missing typesNumbers as strings, wrong date formatState type and format per field
Open enumInconsistent category valuesList allowed values with “one of”

Part 2: Advanced Mistakes (Programmatic Pipelines and Structured Output)

The beginner fixes clean up most single-turn problems. They do not solve the ones that appear when JSON output becomes a contract your code depends on, running thousands of times with no human in the loop. Those require controls outside the prompt text.

The intermediate technique that changes the game: schema-constrained decoding

Before listing the advanced mistakes, it helps to understand the mechanism that removes an entire class of them. Several major providers now offer a mode where you pass a JSON schema (typically in JSON Schema draft format) alongside the prompt, and the model is constrained to emit only tokens that keep the output valid against that schema. OpenAI exposes this as response_format with type: "json_schema" and a strict: true flag. Anthropic exposes a comparable capability through tool definitions where the tool input schema is enforced. Google’s Gemini API offers a controlled generation mode with a responseSchema field.

When you use these modes, the model cannot return invalid JSON — the sampler is not allowed to produce the closing brace early or emit a bare unquoted key. That does not mean your data is correct. It means the syntax is guaranteed and the semantics are still yours to check. A schema can force a field to be a string, but it cannot force that string to be a real email address.

If the option is available to you, use it. It removes the fence problem, the prose problem, and the closed-brace problem in one step. The rest of this section covers what you still have to handle after that.

Mistake 6: Schema drift between prompt and validator

When the schema lives only in the prompt, it can silently disagree with the schema your validator enforces. Someone updates a field name in the prompt and forgets the Pydantic model, or adds a field to the validator that the prompt never mentions. The failure is intermittent and depends on which fields the model fills.

The fix is to make the schema a single source of truth and generate the prompt text from it. Define the schema once — in JSON Schema, Pydantic, or a typed struct — then serialize it into the prompt. This is the one change that eliminates drift by construction.

from pydantic import BaseModel, Field
from typing import Literal

class Ticket(BaseModel):
    order_id: str
    charge_date: str = Field(pattern=r"^\d{4}-\d{2}-\d{2}$")
    card_last_four: str
    issue_type: Literal["double_charge", "late_delivery", "wrong_item", "other"]
    requested_action: str

schema_json = Ticket.model_json_schema()

prompt = f"""Extract the ticket fields. Return only JSON matching this schema:

{schema_json}
"""

Now the schema the model sees and the schema you validate against come from the same object. Change the model, and both update.

Mistake 7: No validation step, so invalid data flows downstream

Even with a perfect prompt, a model will occasionally return a value that is the right type but the wrong content — a date in the wrong format, a category outside the enum, a missing required field. Without a validation step, that data reaches your database or your UI and causes a failure far from where it was created.

Validate every response before using it. Pydantic raises a clear error, and you can decide whether to retry, log, or fall back:

from pydantic import ValidationError

def parse_ticket(raw: str) -> Ticket | None:
    try:
        data = extract_json(raw)
        return Ticket.model_validate(data)
    except (ValueError, ValidationError) as exc:
        # Log the raw response and the error for later inspection
        print(f"Validation failed: {exc}\nRaw: {raw}")
        return None

The value here is not just catching errors early. It is that the validation error tells you what the model got wrong, which is the signal you feed back into a retry.

Mistake 8: Retrying with the same prompt

When validation fails, the instinct is to call the model again with the same prompt and hope for a different result. That works sometimes, but it wastes the information the validator already produced. The error message is a precise description of what was wrong, and you can put it back into the conversation.

A targeted retry appends the validation error and asks for a correction:

def parse_with_retry(raw: str, prompt: str, client, max_attempts: int = 3):
    for attempt in range(max_attempts):
        result = parse_ticket(raw)
        if result is not None:
            return result

        # Send the error back and request only a corrected object
        correction = (
            f"The previous output failed validation with this error:\n"
            f"{last_error}\n"
            f"Return a corrected JSON object only. Do not explain the error."
        )
        response = client.complete(prompt=prompt, messages=[correction])
        raw = response.text

    raise RuntimeError("Exceeded retry budget without valid output")

Two details matter. First, the retry asks for the corrected object and explicitly forbids explanation, which keeps the response parseable. Second, the retry budget is bounded. An unbounded retry loop on a prompt that is fundamentally underspecified will burn tokens and time without converging.

Mistake 9: Assuming structured output means correct output

Schema-constrained decoding guarantees syntax, not truth. A model can return a perfectly valid object with "charge_date": "1901-01-01" because the schema only required a date-shaped string. The syntax passed; the content is wrong.

This is where the advanced techniques end and domain validation begins. For a field like a date or an amount, validate the value, not just the shape. For a field like a category, confirm membership in the enum. For cross-field consistency, check the relationships your schema cannot express — a refund amount that exceeds the original charge, a start date after an end date. Schema constraints handle the container. Business rules handle what goes inside it.

Mistake 10: Over-constraining and forcing the model to hallucinate

There is a real cost to pushing too hard on structure. If you demand a field that the input does not contain, the model has three options: leave it null, omit it, or invent a plausible value. Many prompts discourage nulls, and omission violates a required schema, so the model fills the gap. It fabricates.

A ticket that mentions no order ID but requires one will produce a confabulated number that looks correct and is not. The fix is to make optionality explicit. Mark fields that may be absent, and in the schema allow null. Instruct the model directly: “If a field is not present in the input, use null. Do not guess.” Then make your downstream code handle null rather than assuming presence.

Over-constraining also shows up in enum design. A closed set of five categories that does not include the input’s real category forces the model to shoehorn it into the closest wrong option. Always include an “other” or “unclassified” value.

Advanced Fix Summary

MistakeFailure ModeFix
Schema driftPrompt and validator disagreeGenerate prompt from one schema object
No validationBad data reaches downstreamValidate every response with Pydantic
Blind retrySame error repeats; wasted callsFeed the validation error back, bound attempts
Trusting valid syntaxStructurally valid, factually wrongAdd domain and cross-field checks
Over-constrainingModel hallucinates missing fieldsAllow nulls; include an “other” category

Beginner vs Advanced: When to Use Which

Prompt-level fixes are sufficient when a human reviews the output, the volume is low, and a failed parse costs a copy-paste. In those settings, a clear schema in the prompt and an instruction to omit prose will handle almost everything.

Move to schema-constrained decoding and a validation layer when the output feeds code with no human in between, when the same prompt runs at volume, or when a wrong value has a cost beyond annoyance. The trade-off is real: constrained decoding limits flexibility, adds a schema to maintain, and in some providers can reduce the model’s ability to reason through the response before emitting it. For a task where you want the model to think first and commit to JSON last, forcing structured output from the first token can degrade quality. In those cases a two-step approach works better — reason in free text, then emit JSON in a second constrained call.

Do not add validation and retries to a prompt that is simply missing a schema. The simplest fix is usually the prompt text itself, and the advanced machinery is there to handle the residual failures, not to compensate for an underspecified request.

A Starting Point

If you take one thing from this post, make it the schema. Print the exact JSON shape you expect inside the prompt, generate it from the same object you validate against, and name the types, formats, and enum values explicitly. That single discipline removes the majority of the beginner mistakes. Add validation and targeted retries when you go programmatic, and reserve schema-constrained decoding for the cases where structural guarantees are worth the reduced flexibility. The task is to make the gap between what you asked for and what your parser expects as small as possible — and then to check, rather than assume, that the model stayed inside it.