The cheapest judge is the model you already have running. It is deployed, the client library is already imported, and it is good enough to grade with. So the judge ends up being that model, or its bigger sibling, or the same lab's newest release.
And the thing it grades is usually that same model's output, because that is what you are trying to improve.

If a judge favours its own family, every A-versus-B number you report is tilted before the comparison starts. The tilt is invisible, because what comes out is a quality score and a quality score is what you were expecting to read.

So the question is worth an afternoon. The trouble is that the experiment everybody reaches for cannot answer it, however carefully it is run.
Here is what almost everyone runs first. Ask the model to choose between its own answer and another model's, and see what it says.

Look at both branches. Each one has two explanations and the experiment cannot choose between them.
This is not a sample-size problem. Running it on ten thousand items gives you a very precise number that still means two things at once.

Two things changed at once. That is the oldest mistake in experimental design, and the usual fix applies. Hold the thing you cannot control fixed across both arms, and read the difference rather than the level.
The fix is one extra judge.

Both models answer. Both models judge. The measurement is the difference between the two judges' preference for the same model's text.
This is easier to believe by running it than by reading it. The first half builds data where the truth is known, so you can check what each experiment recovers.
Model A really does write better answers in that simulation, and it really does favour its own by ten points. Ask A alone and you get 71.7%, which is consistent with a large quality gap, with a large bias, or with any mixture. Ask both and subtract, and the answer comes back at +19.6% against a truth of +20.0%. The quality gap cancelled.
The third block is the control, and it is the part that makes this a design rather than a story. Set the bias to zero and the same arithmetic returns -0.2%. A method that only ever returns the answer you are looking for has not been tested.
If one model genuinely writes better answers, both judges lean the same way and the subtraction removes it. What is left cannot be about quality, because quality was held fixed.

Before any of it, send the same prompt to the same judge twice.

Every verdict repeated. So a difference between the two judges is a difference between the judges, not a judge disagreeing with itself.
Worth one caveat, because I saw it. An earlier run of this same lab came back 99 of 100. Judging is stable within a run, and the generated answers are not perfectly stable across runs. That is why the answers in this lab are written once and cached. Without it, a figure built from one run and a terminal capture from another quietly disagree.

Both judges preferred the other model's answers. That is the shared part, and it cancels.
The difference is what the design exists to isolate. It came out at -6 points, which means the author was harder on its own answers than the outside judge was.


The per-item view is the test. On 33 of 50 items the two judges gave identical verdicts, which leaves 17 that can carry any signal at all. Five of those leaned the way self-preference predicts and twelve leaned the other way, which is p = 0.143.
Not shown. And the direction, for what little it is worth, is the opposite of the one being looked for.

The order of that output is deliberate. The power comes before the result, because read afterwards a power calculation is an excuse for a null. The controls come before both.

Step two is the only part that differs from the experiment everybody already runs, and it doubles the cost. Doubling the cost of a test that otherwise cannot answer its question is not an expensive trade.
The test runs on the 17 items where the two judges disagreed. That, and not 50, is the sample size.

A strong lean would have shown up nine times in ten. A mild one would have walked straight past this test.
So the claim is narrow: no strong self-preference between these two models, on these questions. It is not "no self-preference".

The tall block is the reassuring one and it is not the test. The test is whether the two short blocks are lopsided, and five against twelve is not lopsided enough to read.
This design produces a second number for free, and it is the solid one.

Two models from different labs, given the same two answers with no hint about who wrote them, reached the same verdict on 82 of 100 pairs.
That is not a null and it needs no power calculation. It is direct evidence that a judge is reading the answers rather than reading itself, and it is the reassurance the self-preference question was really asking for.
Now the part the design did not fix.

Both models were asked for two or three sentences. One wrote 163.8 words on average and the other wrote 78.

On all 50 items the same model wrote the longer answer. So "picked qwen's answer" and "picked the longer answer" are not two measurements. They are one event with two names.
That does not undo the design. A length preference both judges share cancels exactly as a quality gap does. What it means is narrower: the residue could be a lean toward one's own family, or a difference in how much each judge cares about length. This run cannot tell those apart.

The fix, if you want the cleaner experiment, is to hold length fixed. Truncate both answers to the same word count, or ask for a fixed number of sentences and drop the items where the models ignore it. That is one more afternoon, and this lesson did not spend it, which is why the claim here is the narrow one.

The third row is the one to be strict about. A test that could not have seen a mild effect has not ruled one out, and saying otherwise is the thing this chapter keeps warning about.


Read the first note carefully. The published definition includes the clause "while human annotators consider them of equal quality". That clause is the whole difference between a bias and a model being right, and it is exactly what one judge cannot establish.

The middle branch is usually the answer. A second open model is a download, and a judge with no stake in the answer removes the question instead of measuring it.

The third one is the part I got wrong first. Two runs of this lab at temperature 0 with a fixed seed produced different answers, and the headline moved by an item or two each time. the answers makes the judging reproduce exactly, which is what lets a figure and a terminal capture quote the same numbers.


5 questions - Score 80% to pass
Why can a single judge not measure its own self-preference, however many items you run?
What does taking the difference between two judges remove?
The per-item test came out 5 against 12 on 17 discordant items, p = 0.143. What is the honest report?
One model wrote the longer answer on all 50 items. Why does that matter here?
What is the cheapest way to make the self-preference question go away?
The right-hand column matters and it comes back at the end of this lesson, because this run has one of those and it is not small.
A judge cannot tell you about its own bias. Two judges can, and the second one costs a download.