LLM evaluation is the process of converting a subjective impression of a model response into a repeatable measurement: a defined input set, a defined scoring rule, and a recorded score that two different people would assign the same way. Without that conversion, “the model got better” is a feeling, not a finding.
This tutorial is organized as a myth-versus-reality walkthrough. Each section names a common belief about evaluating AI outputs, contrasts it with what holds up under implementation, and shows the concrete mechanism — code, rubric, or config — that replaces the myth. The final section assembles the pieces into a single pipeline you can run against any prompt change.
Myth 1: “If the output looks good, it is good”
Reality: Visual inspection is the least reliable evaluation method available, and it is the one most teams default to. Human attention drifts, the last example read biases the score of the current one, and there is no record of what “good” meant on Tuesday when you decide on Friday that the new prompt version is worse.
An LLM output is good relative to a task specification, not relative to aesthetics. A response can be fluent, well-structured, and completely wrong. Conversely, a terse, poorly formatted response can be exactly correct. Until you write down what correct means for the task, you cannot distinguish those two cases.
The fix: Define the scoring dimensions before you look at any output. For most text-generation tasks, four dimensions cover the space:
| Dimension | Question it answers | Typical failure it catches |
|---|---|---|
| Correctness | Are the stated facts, numbers, and code references accurate? | Hallucinated APIs, invented statistics |
| Instruction adherence | Did the output satisfy every explicit constraint? | Wrong length, missing required sections |
| Relevance | Does the content address the actual question asked? | Topically adjacent but off-task responses |
| Format compliance | Is the output in the requested structure? | Prose where JSON was required |
Score each dimension on a small integer scale — 0 to 2 is enough for most work. The point is not precision; it is consistency. A 0/1/2 rubric applied the same way across fifty examples tells you more than a 1-to-10 scale applied differently every time.
Myth 2: “A single overall score is enough”
Reality: Collapsing four dimensions into one number destroys the diagnostic value of the evaluation. If a prompt change raises correctness from 1.6 to 1.9 while dropping format compliance from 2.0 to 0.4, the overall score may look flat — and you have lost the information that the change broke JSON output.
Track dimensions separately, then aggregate only for reporting. A weighted sum is fine for a dashboard, but keep the raw per-dimension scores in the dataset. When a regression appears, the dimension breakdown is what tells you where to look.
A useful convention:
{
"example_id": "summarize_042",
"scores": {
"correctness": 2,
"adherence": 1,
"relevance": 2,
"format": 0
},
"weighted_total": 0.45,
"notes": "Summarized the article but included a section header the prompt explicitly said to omit."
}
The notes field is not optional. A score without a one-line justification is unreproducible six weeks later, because nobody remembers why that example got a 1.
Myth 3: “You need a panel of human raters for every example”
Reality: Human labels are the gold standard for calibrating a rubric, but they are the wrong tool for continuous evaluation. A human pass costs seconds to minutes per example; an automated pass costs a fraction of a cent. If your evaluation set has two hundred examples and you change prompts weekly, human review does not scale.
The practical split is:
- Human labels on a small, fixed calibration set — typically 20 to 50 examples. These establish what a 2, 1, and 0 look like on each dimension.
- Automated scoring on the full set, using either deterministic checks or an LLM judge, validated against the calibration set.
The validation step matters. Before trusting an automated scorer, measure its agreement with your human labels. If the automated scorer and the human agree on 90% of a 50-example set, the automated score is a reasonable proxy. If agreement is 60%, the scorer needs rework — usually by making the rubric more explicit.
Cohen’s kappa is the standard agreement statistic for two raters on categorical scores. You do not need to compute it by hand; sklearn.metrics.cohen_kappa_score handles it in one line. A kappa above roughly 0.6 is conventionally treated as substantial agreement for this kind of task.
Myth 4: “Deterministic checks are too simple to be useful”
Reality: Deterministic checks are cheap, fast, and unambiguous, which makes them the correct first layer of any evaluation pipeline. They cannot judge prose quality, but they catch the majority of format and adherence failures before an LLM judge is invoked — and they never hallucinate a verdict.
If your prompt asks for a JSON object with three keys, a json.loads call and a key-presence check tells you whether the output is usable. That check runs in microseconds and costs nothing. Routing every example through an LLM judge when a parser would do is wasteful and introduces avoidable noise.
A runnable starter for the deterministic layer:
import json
import re
def check_json_shape(text: str, required_keys: list[str]) -> dict:
"""Return per-check pass/fail for a JSON-output task."""
results = {}
try:
payload = json.loads(text)
results["valid_json"] = True
except json.JSONDecodeError:
return {"valid_json": False}
results["has_required_keys"] = all(k in payload for k in required_keys)
return results
def check_constraints(text: str, max_words: int, banned_phrases: list[str]) -> dict:
"""Deterministic adherence checks for free-text tasks."""
word_count = len(re.findall(r"\b\w+\b", text))
lowered = text.lower()
return {
"under_word_limit": word_count <= max_words,
"word_count": word_count,
"no_banned_phrases": not any(p.lower() in lowered for p in banned_phrases),
}
sample = '{"summary": "Short answer.", "confidence": 0.8}'
print(check_json_shape(sample, required_keys=["summary", "confidence"]))
print(check_constraints(sample, max_words=100, banned_phrases=["as an AI"]))
Run this against a batch of model outputs and you immediately know how many failed on structure alone. In practice, a meaningful fraction of prompt-iteration regressions are structural — a trailing sentence after the JSON, a missing key, a word-count overflow — and this layer catches them without any model call.
Myth 5: “An LLM judge is just the model grading itself, so it’s circular”
Reality: An LLM judge is circular only if it is the same model, with the same prompt, evaluating its own output without a rubric. The whole technique depends on breaking all three of those conditions.
Three controls make an LLM judge useful:
- Different model family. If the generator is one vendor’s model, use a different vendor’s model as the judge. Cross-family judging reduces shared blind spots.
- Explicit rubric with anchors. The judge prompt must define what a 2, a 1, and a 0 mean for each dimension, with example phrases for each level. Vague instructions like “rate the quality” produce unusable scores.
- Score-then-justify order. Require the judge to output its reasoning before the numeric score. If the score comes first, the justification becomes post-hoc rationalization.
A judge prompt that satisfies all three:
You are evaluating a model response against a task specification.
Task specification:
{task_spec}
Model response:
{model_response}
Score the response on each dimension using this rubric:
correctness:
2 = every factual claim is accurate and verifiable from the task specification
1 = contains one minor inaccuracy that does not change the overall answer
0 = contains a factual error that invalidates the answer
adherence:
2 = satisfies every explicit constraint in the specification
1 = misses exactly one constraint
0 = misses two or more constraints, or ignores the core instruction
format:
2 = matches the requested structure exactly
1 = minor structural deviation (extra whitespace, alternate key order)
0 = wrong structure (prose instead of JSON, missing required sections)
For each dimension, first write one or two sentences of justification,
then output the numeric score on its own line as "dimension: score".
Parse the output with a regex that captures the dimension: score lines. Reject any response where a dimension is missing — a judge that skips a dimension is a judge that is not following the rubric, and its remaining scores are suspect.
Myth 6: “Absolute scores from different runs are comparable”
Reality: LLM judges drift. The same judge model, given the same response on two different days, will not always return the same score — temperature, prompt wording, and model updates all introduce variance. Comparing absolute scores across runs without accounting for that variance produces false conclusions about whether a prompt change helped.
The robust alternative is pairwise comparison. Instead of scoring each response on an absolute scale, present the judge with two responses — one from the baseline prompt, one from the candidate — and ask which is better on a given dimension. Pairwise judgments are more stable because the judge only has to compare, not calibrate.
Run the comparison twice with the order swapped. If the judge prefers A in the first pass and A again in the second (now positioned second), the preference is order-independent and can be trusted. If the judge flips when the order flips, that example is a tie for practical purposes — record it as such rather than forcing a winner. This position-bias check is the single most important control for pairwise judging, and it is the one most often skipped.
import anthropic
client = anthropic.Anthropic()
def pairwise_judge(task: str, response_a: str, response_b: str) -> str:
"""Return 'A', 'B', or 'tie'. Run twice with swapped args to check position bias."""
prompt = (
f"Task: {task}\n\n"
f"Response A:\n{response_a}\n\n"
f"Response B:\n{response_b}\n\n"
"Which response better satisfies the task? "
"Answer with exactly one token: A, B, or tie."
)
msg = client.messages.create(
model="claude-sonnet-4-5",
max_tokens=4,
messages=[{"role": "user", "content": prompt}],
)
return msg.content[0].text.strip()
def run_with_bias_check(task, resp_new, resp_old):
forward = pairwise_judge(task, resp_new, resp_old) # A=new
reverse = pairwise_judge(task, resp_old, resp_new) # A=old
# Normalize reverse back to new/old naming
normalized = {"A": "B", "B": "A", "tie": "tie"}[reverse]
if forward != normalized:
return "tie" # position bias detected
return forward # 'A' means new wins, 'B' means old wins
The two-call pattern doubles the judge cost per example. That is the price of a trustworthy preference signal, and it is cheaper than acting on a biased one.
Myth 7: “One good evaluation run is enough”
Reality: A prompt that wins on 50 examples may lose on the next 50. Evaluation sets overfit the same way models do — after enough iterations against the same set, the prompt has been tuned to that set rather than to the task.
The mitigation is a held-out split. Before you start iterating, set aside 20 to 30 percent of your examples and do not look at them during development. Run them only when you believe the work is finished. If the candidate prompt’s advantage evaporates on the held-out set, the improvement was an artifact of overfitting and the prompt has not generalized.
This is the same discipline as a train/test split in machine learning, applied to prompt engineering. The cost is that you have fewer examples to iterate against, which slows iteration. The benefit is that you find out whether your improvement is real before you ship it, rather than after. Use the held-out set sparingly — every peek erodes its value.
Myth 8: “Evaluation is a one-time setup task”
Reality: Evaluation is a maintenance burden that grows with the product. Three forces push it out of date:
- Prompt drift. Prompts get edited incrementally. Each edit invalidates the previous evaluation baseline unless the set is re-run.
- Model updates. Vendors ship new model versions with changed behavior. A prompt that scored 1.9 last quarter can score 1.6 after a silent update.
- Task shift. The inputs users send change over time. An evaluation set built from March traffic may not resemble June traffic.
The practical response is to version the evaluation set alongside the prompt, and to re-run the full set on every model upgrade. Treat the evaluation set as a test suite, not a report. If a prompt change or a model swap drops the aggregate score by more than a defined threshold — commonly two to three percentage points on a weighted total — the change does not ship.
Step-by-Step: A Concrete Pipeline
The sections above are principles. Here is the same material as an ordered procedure you can implement this week.
Step 1 — Build the evaluation set. Collect 50 to 100 input examples that represent the real distribution of user requests, not the easy cases. Each example is an object with an ID, the input, and any reference output if one exists.
Step 2 — Write the rubric. For each dimension, define what 0, 1, and 2 mean, with anchors specific to your task. Store the rubric as a string constant so it can be injected into both judge prompts and human labeling instructions.
Step 3 — Label a calibration subset. Hand-score 20 examples on every dimension. These are the ground truth against which automated scorers will be validated.
Step 4 — Validate the automated scorers. Run the deterministic checks and the LLM judge on the calibration subset. Compute agreement against the human labels. If kappa is below roughly 0.6, tighten the rubric and repeat.
Step 5 — Run the full set. Score all examples on both the baseline prompt and the candidate prompt. For each, record per-dimension scores.
Step 6 — Compare. Aggregate per-dimension averages and run the pairwise, position-bias-controlled comparison. Report both signals, because they sometimes disagree — a candidate can win on average absolute score while losing in head-to-head comparisons, which usually means the absolute scores are being driven by a few outliers.
Step 7 — Verify on the held-out set. Run the winning prompt against the reserved examples. Only promote to production if the advantage holds.
When NOT to Use This Approach
The pipeline above is overkill for three common cases.
One-off generations. If you need a single response to a single question, evaluating it formally is more expensive than rereading it and deciding. Evaluation pays off only when you will generate the same task repeatedly.
Highly subjective tasks. Creative writing, brand voice, and humor resist rubric definition. Pairwise human comparison is still useful, but automated judges on subjective style dimensions tend to produce noise. For these, keep the human in the loop and accept lower throughput.
Tasks where correctness is undefinable in advance. Open-ended research questions have no reference answer. Deterministic checks catch structural failures, but scoring content quality requires domain expertise that no judge prompt supplies.
The Minimum Viable Evaluation Setup
If you take nothing else from this post, take this: a small fixed set, a written rubric, and pairwise comparison with an order-swap check. Those three elements cost a few hours to build and catch the majority of prompt regressions that otherwise ship silently. The deterministic layer handles structure. The LLM judge handles content quality. The held-out split tells you whether the improvement is real.
Everything beyond that — multi-judge ensembles, human rater pools, statistical significance testing — is refinement on top of a foundation that most teams have not built yet. Start with the foundation.
The single most common evaluation mistake is not choosing the wrong judge or the wrong metric. It is not evaluating at all, because “it looked fine when I skimmed it.” Every prompt change that ships without a measurement is a coin flip dressed up as a decision.
🔗 Recommended Reading
- LLM Guardrails for Beginners: A Step-by-Step Tutorial to Filtering Unsafe Outputs
- Building Your First AI Agent: A Beginner Tutorial with Python and the OpenAI API
- Function Calling in LLM APIs: A Beginner Tutorial for Connecting AI to Real Tools
- RAG Chunking Strategies for Beginners: A Step-by-Step Tutorial
- Prompt Engineering Mistakes That Break JSON Output and How to Fix Them