LLM Evaluation: Methods, Judge Bias, and Shrinking the Surface
Every guide to LLM evaluation recommends using a model as a judge. Almost none of them cite the research on how those judges fail. Here is what the papers actually found, and why the size of your evaluation problem is an architectural choice.

What the engineering leader actually needs to know
Five of the most substantial guides to LLM evaluation ranking today carry dedicated sections on using a model as a judge. Between them they cite no research on whether model judges work. SuperAnnotate's guide has two headings promising exactly that answer, "How Do I know if My AI Judge is Working Well?" and "When Does LLM-as-a-judge Struggle?", and its only arXiv link is to a paper about model collapse from synthetic data (https://arxiv.org/abs/2307.01850). Arize's page, the freshest serious competitor on the query, runs an LLM-as-a-judge section and an agent-as-a-judge section, cites tau-bench and GAIA, and cites nothing at all on judge reliability.
The literature exists. It is public, it has numbers in it, and it is more useful to you than any framework comparison.
Evaluation, as I am using it here, means measuring whether a model-backed application produces acceptable outputs and acceptable actions across a representative set of inputs, with a stated method for each property you care about. That last clause is where most programs fail. They pick a metric before they pick a method, and end up with one score covering properties that need three different instruments.
For the person who has to sign off on expanding an AI feature, evaluation answers two questions. Does it work on inputs that look like production, and would anyone notice if it stopped. Everything else is instrumentation for those two.
Model evals, application evals, and agent evals are three different jobs
Most confusion about this topic comes from collapsing three jobs into one word.
A model eval scores a model against a public dataset. MMLU, HumanEval, GPQA. It tells you something real about the model and almost nothing about your application, because your application has a prompt, a retrieval layer, a tool inventory, and a user population that no public benchmark ever saw.
An application eval scores your system on your inputs. An agent eval scores the path the system took to get there.
| Layer | What it answers | What it cannot answer | | --- | --- | --- | | Model eval (benchmarks) | How this model compares to others on public tasks | Whether your prompt, retrieval, and tools work on your traffic | | Application eval | Whether your system produces acceptable outputs on your inputs | Whether the intermediate steps were sound, or repeatable | | Agent eval | Whether the trajectory was sound: right tools, right order, recovery after an error, consistency across repeated runs | Whether the final answer was useful to the person who asked, which still needs output-level scoring |
Benchmarks are worth reading. They are the reasonable starting point for routing between models before you have data of your own. They are not evidence your feature works.
The four ways to evaluate, and what each is actually for
Code-based checks
The cheapest and most reliable method, and the one teams skip. Schema validity, required fields present, every citation URL resolving, p95 latency inside budget, no PII patterns in the output, no illegal state transition on a write. These are predicates. They pass or they fail, with no rubric and no agreement rate to maintain.
Teams skip this layer because the failures that worry them feel semantic, then discover half their incidents were malformed JSON and a dead link. If a property can be written as a predicate over the output, write the predicate. Never send it to a judge.
Model-based judges
For the properties that resist assertions. Groundedness against retrieved evidence, relevance to the question asked, tone against a rubric, whether a refusal was appropriate. A judge is a model handed an output, optionally a reference, and a rubric, and asked to score. It works, under conditions the next section covers.
Human review
Calibration and high-risk decisions. Human labels are the thing a judge's agreement rate gets measured against, so a program with no human labels has no way to know whether its judge has drifted. Human review does not scale, and human reviewers disagree with each other, which is why the MT-Bench work reports human-to-human agreement as a ceiling rather than as ground truth. Budget it as a sampling instrument, not as a gate on every release.
Production signals
A curated dataset contains the failures you already thought of. Production contains the rest. Thumbs, escalation rate, retry rate, tool error rate, abandonment, and the trajectory logs that let you reconstruct a bad run after the fact. This layer depends entirely on observing what the agent actually did, which is a separate engineering investment with its own cost.
| Method | What it measures well | What it cannot measure | Relative cost | | --- | --- | --- | --- | | Code-based checks | Schema validity, required fields, resolvable citation URLs, latency budgets, forbidden content patterns, legal state transitions | Whether text is grounded, relevant, or appropriately toned | Lowest. Runs in CI with no model call | | Model-based judges | Groundedness, relevance, tone, refusal appropriateness, pairwise preference | Its own error rate, and anything it is biased on without correction | Moderate per run, plus the standing cost of measuring agreement | | Human review | Calibration labels, high-risk decisions, failure modes nobody anticipated | Volume. Reviewers also disagree with each other | Highest per label | | Production signals | Failures absent from your dataset, drift over time, consistency across repeated runs | Anything not instrumented, and root cause without trajectory logs | Low marginal, high setup |
What the research says about LLM judges
Model judges work. Zheng et al., "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena" (https://arxiv.org/abs/2306.05685), reports GPT-4 judges reaching over 80% agreement with human preferences, which matches the human-to-human agreement level on the same comparisons. Any argument that model judges are unusable has to account for that number, and none of them do.
The same paper names the failure modes: position, verbosity, and self-enhancement bias. Three later papers isolated them with numbers.
Position bias is the cheapest to defend against and the most alarming when you see it. Wang et al., "Large Language Models are not Fair Evaluators" (https://arxiv.org/abs/2305.17926), found that swapping the order of two candidate responses was enough to move the verdict. With ChatGPT as the evaluator, one ordering gave Vicuna-13B wins over ChatGPT on 66 of 80 questions.
Verbosity bias has a published correction. Dubois et al., "Length-Controlled AlpacaEval" (https://arxiv.org/abs/2404.04475), applies a regression-based adjustment that estimates preference as if both responses were the same length, and reports correlation with Chatbot Arena rising from 0.94 to 0.98.
Self-preference has the strongest causal evidence. Panickssery et al., "LLM Evaluators Recognize and Favor Their Own Generations" (https://arxiv.org/abs/2404.13076), found GPT-4 and Llama 2 identify their own outputs at above-chance rates, with a linear relationship between self-recognition ability and self-preference, and controlled experiments supporting a causal reading. Liu et al., "G-Eval" (https://arxiv.org/abs/2303.16634), had warned about LLM evaluators favoring LLM-written text a year earlier, and reported Spearman 0.514 with human judgments on summarization. That is a genuine correlation, and it is a long way from interchangeable with a human.
| Bias | What the research found | Practical mitigation | | --- | --- | --- | | Position | Wang et al. (arXiv 2305.17926): with ChatGPT as evaluator, reordering candidates alone produced Vicuna-13B wins on 66 of 80 questions | Score both orderings and treat a disagreement as a tie | | Verbosity | Dubois et al. (arXiv 2404.04475): length control raised correlation with Chatbot Arena from 0.94 to 0.98 | Apply length control, or normalize response length before scoring | | Self-preference | Panickssery et al. (arXiv 2404.13076): above-chance self-recognition, with a linear relation between self-recognition and self-preference | Judge with a different model family. Never let a model grade its own generation |
These papers date from 2023 and 2024, and the models in them are two or three generations behind what you are running. Read them as documented failure modes to test for on your judge and your rubric, not as fixed properties of current models. The test is cheap: reorder your candidate pairs, measure the flip rate, and you have your own position-bias number by the end of the afternoon.
Which leads to the position I will defend. A judge score with no measured agreement rate against human labels is not a metric. It is a vibe with a decimal point on it. Report agreement alongside any judge-derived number, or do not report the number.
RAG and agent evaluation
Retrieval and generation fail separately, so measure them separately. Retrieval quality asks whether the answer was in the retrieved set and at what rank. Generation quality asks whether the output stayed inside the evidence it was given. A fluent, confident answer built over the wrong documents scores well on every output-level metric and is still wrong, and you will not find that by scoring the output alone.
For agents, final-answer scoring hides mid-trajectory failure, and it hides inconsistency completely. Yao et al., "tau-bench" (https://arxiv.org/abs/2406.12045), makes this point better than any vendor page. It scores the end state of the database rather than the chat transcript, and it introduces pass^k, which asks whether an agent succeeds on the same task across repeated trials. In the retail domain, pass^8 fell below 25%. Agents that pass once frequently do not pass eight times. Mialon et al., "GAIA" (https://arxiv.org/abs/2311.12983), reports humans at 92% against GPT-4 with plugins at 15% on questions humans find conceptually simple.
Consistency is a different property from correctness, and it is the one that decides whether a workflow can run unattended.
What good looks like in practice
Take a support-reply drafting agent grounded in a knowledge base. The design work is a split, made property by property, before any tooling decision.
In CI on every change, as code-based assertions: the response matches schema, required fields are present, every citation URL resolves and points to a document that was actually in the retrieved set, no PII pattern appears in the draft, p95 latency is inside budget, and the ticket status transition is legal.
On a versioned dataset, with a model judge: groundedness against the retrieved passages, relevance to the customer's question, tone against the rubric.
Sampled weekly, by a person: escalation decisions, plus a labeled slice used for nothing except measuring judge-human agreement.
The program around that split:
- Write the failure list before the metric list. Ask what a bad run looks like to the customer, then work backwards to what would catch it.
- Assign every property to exactly one method. If two methods cover it, keep the cheaper one.
- Build the dataset from real traffic, version it, and grow it from incidents. Every production failure becomes a test case.
- Put the code-based checks in CI and block merges on them. Judges are too slow and too noisy to gate a pipeline.
- Measure your judge before you trust it. Agreement against human labels, flip rate under reordering, and a check that it is not scoring its own family's output.
- Sample production continuously, and keep enough trajectory detail to reconstruct a specific bad run months later.
- Define in advance what a moved score triggers. A number that changes and produces no decision is not being used.
NIST's AI Risk Management Framework (NIST AI 100-1, released 26 January 2023) files all of this under its Measure function, and the Generative AI Profile (NIST AI 600-1, 26 July 2024) extends it to generative systems. Neither document prescribes your metrics. Both expect you to state a method, document it, and be able to show the measurement was taken. That is a low bar, and most evaluation programs miss it, because the method lives in a notebook on one engineer's laptop. If your program exists to satisfy an audit, the artifact is the documented method plus the retained record, which is the same problem as governance from policy to implementation and control at the point of action.
Where vendors are getting this wrong today
Four patterns, all checkable on the current results page for this query.
Judge sections with no citations, named at the top of this piece. Two of the strongest pages on the query promise judge-reliability answers and supply none.
Comparisons locked in images. SuperAnnotate's method comparisons are pictures, which means no crawler and no answer engine can read them, and neither can a screen reader.
Governance sections that skip NIST. Arize's regulatory section cites the EU AI Act, SR 11-7, and the FDA, and omits NIST entirely, which is the most-referenced US framework in this area and free to cite.
Self-ranking tool lists. SuperAnnotate's "Top 10 LLM Evaluation Frameworks and Tools" puts SuperAnnotate first.
One more, offered as a caution rather than a criticism: NVIDIA's evaluation post is the most citation-dense page in the set, and two of its arXiv labels are mismatched. If you are building a reading list from a vendor page, open the papers.
The question the metric debate skips
Every page on this query treats the evaluation surface as fixed and competes on how to cover it. The surface is not fixed.
What needs evaluating is a consequence of design. A step that re-reasons on every run can fail differently tomorrow than it failed today, so it needs evaluating for as long as it exists, with judges, rubrics, agreement rates, and drift monitoring behind it. A step that runs as code fails the same way every time, and an assertion catches it.
So there is a question that comes before metric selection. How much of this workflow is probabilistic, and did anyone actually decide that. Extracting an invoice number. Checking a contract term against a table. Writing a status back to the system of record. In most agent architectures these get re-reasoned on every single run because that is the default, and each one then joins the evaluation backlog permanently. How the workflow is structured sets how much of it you will be evaluating for the rest of its life.
What we're doing about this at Major
Our position is that evaluation is treated as a coverage problem when a large part of it is a scoping problem. Teams inherit a probabilistic surface, accept it as the shape of the work, and then spend engineering quarters instrumenting it. The surface was a design decision that nobody made on purpose.
On Major, when an agent works out how to handle a repeatable part of a task, it builds an app for that part and runs the app instead of reasoning through the step again. That app is deterministic code with its own managed database, storage, and logs. You evaluate it the way you evaluate any code, with assertions that pass or fail, and there is no judge whose agreement rate you also have to track. The evaluation effort then concentrates on the judgment calls, which are the steps that needed a model in the first place.
This does not make evaluation unnecessary, and it does not take the model out of the system. The judgment calls still need real evaluation, judges included, with every bias caveat above still applying to them. What changes is how many steps are in that category, and whether the number grows every time you ship. Reason once. Run forever.
If your evaluation backlog keeps growing because every step in the workflow re-reasons, the fix sits upstream of your eval suite. You can see how Major's agents move repeatable steps into deployed apps you test with assertions rather than scores.
Related articles
- AI Agent Governance: Why Control Must Move to the Point of Action
- How to Build an Agentic Workflow: Patterns, Steps, and Limits
- LLM Router: Routing Strategies and Their Real Limits
Related articles
Frequently asked questions
- How do you evaluate LLM response quality?
- Split the output's properties by method. Code-based assertions handle the deterministic ones: schema, required fields, resolvable citation links, latency. A model judge scores semantic properties like groundedness, relevance, and tone. Human labels calibrate the judge and cover high-risk decisions. The rule for choosing between them: if a property can be written as a predicate that passes or fails, never send it to a judge.
- What is the difference between LLM evals and benchmarks?
- Benchmarks score a model against public datasets like MMLU or HumanEval, which tells you about the model. Evals score your application against your data, your prompts, your retrieval layer, and your requirements. A model that tops a benchmark can still fail your traffic, because no public dataset contains your customers' questions. Benchmarks inform model selection. Evals decide whether your feature ships.
- Is LLM-as-a-judge reliable?
- Reliable enough to use, not reliable enough to trust unmeasured. Zheng et al. (https://arxiv.org/abs/2306.05685) report GPT-4 judges reaching over 80% agreement with human preferences, matching human-to-human agreement. The same line of work documents position bias (https://arxiv.org/abs/2305.17926), verbosity bias (https://arxiv.org/abs/2404.04475), and self-preference (https://arxiv.org/abs/2404.13076). Randomize candidate order, control for length, avoid self-judging, and report agreement against human labels.
- What are rubrics in LLM evaluation?
- A rubric is the explicit scoring criteria handed to a judge or a human reviewer: what each score means, what counts as grounded, what counts as an appropriate refusal. Judge scores are sensitive to how a rubric is worded, so treat it as a versioned artifact. Changing the rubric changes your numbers, so version it alongside the dataset.
- What are the best LLM evaluation tools?
- Judge by criteria rather than by rankings. Does it run code-based assertions as well as model judges, can it measure judge-human agreement, does it version datasets, does it run in CI, and can you export the records an auditor will ask for. The tool matters less than whether your program measures its own judges. Major goes further upstream, letting agents move repeatable steps into deployed apps so those steps get assertions instead of scores.