Essay
Before You Trust an LLM Judge, Check Whether the Judges Agree
LLMs are increasingly used to evaluate the outputs of other LLMs. They compare responses, assign scores, assess properties such as safety or fairness, and increasingly replace at least part of the human evaluation process. This makes evaluation faster and…
LLMs are increasingly used to evaluate the outputs of other LLMs. They compare responses, assign scores, assess properties such as safety or fairness, and increasingly replace at least part of the human evaluation process. This makes evaluation faster and easier to scale, but it also introduces another model into the measurement process.
Before relying on that model, there are some basic questions worth asking.
Do different LLM judges agree? Does that agreement hold when the context becomes more consequential? Is the judge independent of the model that generated the output? And can we inspect why the judge gave the score that it did?
These questions motivated our new paper in Technological Forecasting and Social Change, in which Joni Salminen, Bernard J. Jansen, and I examined inter-rater agreement among GPT-4o, Claude 3 Sonnet, and Gemini 1.5 Pro when evaluating AI-generated personas. The judges assessed fairness, lack of stereotypicality, and diversity across 60 individual personas and 12 persona sets generated by four LLMs.
The main result is not that LLM judges fail. At the individual-persona level, agreement was moderate overall. The more important result is that this agreement was not constant. Depending on which model generated the personas, agreement ranged from .39 to .63. Across the educational contexts we studied, individual-level agreement declined from ICC = .55 for peer tutoring to .34 for mandatory training. The three scenarios differed in more than risk, so we cannot attribute that decline to risk alone. But the result does show that reliability observed in one setting should not automatically be assumed in another.
This matters because LLM judges are often treated as though reliability belongs to the model itself: if a judge works reasonably well on one benchmark, it is reused elsewhere. Our results suggest a different interpretation. Reliability belongs to the evaluation procedure under particular conditions, including what is being judged, how it is scored, and the context in which the evaluation takes place.
Agreement should be checked before scores are trusted
Using one LLM judge gives us a score. Using several judges lets us ask whether that score depends heavily on which model happened to produce it.
Agreement is not the same as accuracy. Three models can agree and still share the same error. Our study does not compare the LLM judges with human experts, so it cannot establish whether their judgments were correct.
But disagreement still tells us something useful. If several judges receive the same material, rubric, and instructions and produce substantially different evaluations, treating one of those scores as an objective measurement becomes difficult to justify.
This is why I would treat inter-rater agreement as an early diagnostic rather than as a final validation step. Before asking whether an LLM judge is right, first establish whether another plausible judge reaches a reasonably similar assessment.
Anthropic recently made a related point in its guidance for rigorous wellbeing evaluations. It recommends calibrating graders against human expert labels, reporting agreement, inspecting evaluation transcripts manually, and measuring reliability and reproducibility. It also lists the absence of human validation or reported inter-rater agreement as a common weakness in evaluations. The guidance concerns wellbeing evaluations specifically, but the measurement principle applies more broadly: the grader itself needs to be evaluated before its outputs are treated as evidence.
Ask the judge to justify the score
Agreement coefficients alone are not enough.
For each judgment, I would also require a short written justification tied directly to the rubric: one or two sentences explaining what in the evaluated output supports the assigned score.
The purpose is not to treat this explanation as a faithful account of the model’s internal reasoning. It is to create an auditable part of the evaluation. If a model assigns a low fairness score, for example, the evaluator should be able to inspect the accompanying sentence and ask whether it identifies something relevant to the stated fairness criterion.
This creates a simple sense check at each stage:
- Score: What did the judge decide?
- Justification: What evidence does it give for that decision?
- Rubric check: Does that evidence actually correspond to the criterion being scored?
- Agreement check: Do other judges reach a similar conclusion?
- Human check: For a sample, do human reviewers consider the judgment and justification defensible?
This is particularly useful when numerical agreement appears acceptable but the models are interpreting the rubric differently.
We saw a related problem in our own results. The protocol required textual support for the judges’ ratings. Non-compliance ranged from 9.6% to 26.7% at the individual-persona level, but increased to 34–49% when the judges evaluated complete persona sets. A numerical score by itself would hide that failure to complete the requested evaluation.
Anthropic’s guidance again points in a similar direction. It recommends clearly defined constructs, decomposed scoring, and manual inspection of evaluation transcripts. A short rubric-linked justification provides something concrete to inspect during that process.
Check whether reliability changes with risk or context
The second question is whether an agreement result transfers to the setting in which the judge will actually be used.
Our individual-level agreement declined from .55 in peer tutoring to .34 in mandatory training. We cannot say that increased risk caused the difference, because these scenarios were not identical apart from risk. The result is still useful because it demonstrates that the same judges did not maintain the same level of agreement across contexts.
This is particularly relevant for higher-stakes applications. Evidence that a judge performs consistently on ordinary or low-risk examples should not be used as sufficient evidence that it will remain reliable when the evaluation becomes more difficult or consequential.
Anthropic recommends explicitly covering multiple severity levels and reporting results separately by severity. Our results support the broader logic of that recommendation. Evaluations can conceal important differences when scores or reliability estimates are pooled across contexts.
The practical implication is straightforward: measure agreement at the levels of risk or context about which you intend to make claims.
Avoid circular evaluation where possible
There is also a structural problem with LLM-based evaluation.
The models producing the outputs and the models evaluating them are often the same models or belong to the same model families.
Our study contains exactly this overlap. GPT-4o, Claude 3 Sonnet, and Gemini 1.5 Pro appeared both among the persona generators and among the judges. Our analysis does not isolate whether this produced self-preference, so we cannot claim from this study that a model favored its own personas.
But the possibility matters. Previous work has found evidence that LLM evaluators can recognize and favor their own generations. More generally, asking a model to generate an output and then asking the same model to decide whether that output is good weakens the independence of the evaluation.
Anthropic identifies this explicitly as “circular AI-as-judge”, meaning that a model grades itself or its relatives. Its guidance recommends considering a different LLM for scoring than the models being evaluated.
Where possible, I would therefore separate the generator and judge panels and use judges from different model families. This does not guarantee independence, since models may still share training data, conventions, or biases, but it removes the most direct form of circularity.
A practical starting point
Putting these pieces together, I would start an LLM-based evaluation with a fairly simple procedure.
Use more than one judge. Ask each judge for a score and a brief rubric-linked justification. Check whether that justification actually supports the score. Measure agreement among the judges rather than reporting only their average. Break the agreement results down by relevant contexts or severity levels. Keep the judges separate from the models generating the content where possible. Inspect a sample manually, and calibrate against human expert judgments when the evaluation is consequential enough to require it.
This is close to the direction outlined in Anthropic’s evaluation guidance: clearly define the construct, decompose the scoring, test different levels of severity, validate graders, inspect transcripts, measure reliability, and avoid circular AI-as-judge.
Our paper provides empirical evidence for why several of those checks matter. The same three judges did not agree equally across generators and contexts. Their ability to complete the requested evaluation also deteriorated substantially when moving from individual personas to persona sets. And the experimental design itself exposes the circularity problem created when the same model families appear on both sides of an evaluation.
The conclusion is therefore not that LLM judges should be avoided. It is that an LLM-generated score should not become a measurement simply because it is easy to produce.
Before trusting the score, I would ask four questions:
Do different judges agree?
Does that agreement hold in the context and level of risk that matters?
Can I inspect a short justification and determine whether the score follows from the rubric?
And is the judge sufficiently independent of the model whose output it is evaluating?
Those questions do not solve LLM evaluation. But they make it considerably harder for an unreliable evaluation procedure to look reliable simply because it produces a number.
Paper:
Amin, D., Salminen, J., & Jansen, B. J. (2027). Assessing inter-rater agreement among large language model judges: Evidence from persona evaluation across varying educational risk contexts. Technological Forecasting and Social Change, 234, 124894.
https://doi.org/10.1016/j.techfore.2026.124894