A rubric is the sentence in a judge prompt that says what to look for. This one was vague, so somebody made it clearer. The pull request reads better than what it replaced and the review approved it in a minute.
Next week the pass rate has moved and the team goes looking for a bug in the model.

An earlier lesson in this chapter shows a rubric-sensitivity chart and marks it, correctly, as illustrative: five phrasings, a twenty-point band, numbers chosen to make the shape visible. This lesson measures the real thing, and asks a second question that chart could not.

The size of the band is the part an illustration can carry. What it cannot carry is whether six phrasings disagree only about the level, or also about which answers are better. That second question decides what you can still compare after a reword, and it is the one this lesson was built to answer.
Six sentences. Nothing else in the prompt changed: same question, same answer, same instruction to reply PASS or FAIL, same judge at temperature zero.

None of those is wrong. None is a trick. Each one is something you would find in a real repository, written by somebody trying to be clear.

The control comes first, as always. Every verdict repeated, so a difference between two phrasings is a difference between the phrasings.

85.0% to 96.7%. The answers never changed.
If your release gate says "85% must pass", that gate is a statement about a sentence in your rubric as much as about your system. Under one wording you are at the line. Under another you are ten points clear of it.
One more thing about that illustrative chart. It puts "grade strictly for accuracy" at the harsh end and "rate the quality" at the lenient end, which is the intuition most people have. This judge does the opposite. "Rate the quality" was the harshest of the six at 85.0%, and "strictly accurate" was second harshest at 88.3%. The size of the band was about right. The direction was not, and it is not something you can reason your way to from the sentences. You have to run it on your judge and your answers.
That 11.7 points is the number a dashboard shows, and it is the wrong one to quote.

Most of the set is answers that every phrasing passes. They sit in the denominator and cannot move, so they drag the measured spread toward zero.

Eight answers out of sixty carry the entire effect. That is the same arithmetic as the effective-sample-size lesson, arriving from a different direction.

Quote the 35%. It describes what a reworded sentence does to the answers a judge has to think about. The 11.7% describes what it does to a file that is mostly answers nobody disagrees about, and a golden set fills up with those.

Section 2b is the one that says how big the problem is. Section 3 is the one that decides what you can still trust.
Here is the question the illustrative chart could not ask. When the wording moves the pass rate, is it moving the level or the order?

That distinction decides what breaks and what survives, and it is not visible in a chart of pass rates.

Fourteen of fifteen pairs were nested: one phrasing's passes sat entirely inside the other's. The phrasings almost always agree about which answers are better, and differ about how many clear the bar.
The pass sets are the artefact worth keeping, and once you have them the rest is set arithmetic.
Part 3 refines the headline. Restricted to the 20 weak answers, all 15 pairs are nested: on the answers a judge has to forgive or not, the wordings never disagree about the order. The single crossing pair is carried by one good answer that "strictly accurate" failed and "rate the quality" passed, which is a smaller and different problem.
Part 4 is the practical one. A gate of "90% must pass" is cleared by four of the six wordings and missed by two, on the same system and the same answers. Part 5 reads the very same verdicts as a comparison instead. Each clear pair holds a better and a worse answer to one question. So the worse answers and the better answers are two systems. Every phrasing really did grade both, side by side. All six put the better system ahead.

A threshold does not survive a reword. A comparison does, as long as both sides are graded by the same sentence.

The first question takes a git log and it is the one people skip, because a rubric edit does not look like a change to the measurement. It is the measurement.

Nobody in that sequence did anything wrong. The rubric edit was an improvement, the eval ran exactly as written, and the dashboard reported what it measured. The day goes into the model because a rubric is not on anybody's list of things that count as the measurement.

Asked one at a time, this judge passes most of the deliberately weak answers under most wordings. Asked to choose between the good answer and the weak one of the same pair, the same judge picks the better one almost every time. An earlier lesson in this chapter measured that directly: three wrong out of forty.
Pointwise it forgives. Pairwise it discriminates. Same model, same answers, a different question about them.

Those blocks climb in a line. That shape is what a moving threshold looks like. A changing ranking would not be monotone.

Two things in that abstract are worth carrying. A bigger judge does not make the sensitivity go away, so reaching for a stronger model is not the fix. And the recommendation is a habit rather than a technique: report a range across plausible formats instead of a single number.


Adding an example to a rubric feels like documentation. Reordering the criteria feels like tidying. Both are edits to a measuring instrument.

Read those two sentences and decide which one you meant. That decision is the eval, and it is usually made by accident.


Storing pass sets rather than counts is the detail worth copying. Counts tell you the level moved. Only the sets tell you whether the order did, and that is the question that decides what you can still compare.

5 questions - Score 80% to pass
Six rubric phrasings moved the headline pass rate 11.7 points but moved the weak-answer pass rate 35 points. Why is the headline the wrong number to quote?
Fourteen of fifteen phrasing pairs were nested. What does nested mean here, and why does it matter?
Your pass rate dropped 8 points this week. What is the first thing to check?
The judge passed between 11 and 18 of the 20 deliberately weak answers, but picks the better answer almost every time when asked to compare the two answers of the same pair. What does that suggest?
What does the published work on prompt sensitivity recommend reporting?
Read the margin column as well, because it is the part that is easy to miss. Under the harshest wording the better system is 9 answers ahead out of 20. Under the most lenient it is 2 ahead. The direction survived the reword. The size of the gap did not. A lenient rubric leaves you a true comparison you may not have the sample size to see. That is the eval-power lesson again, reached from the rubric rather than from the item count.

The contested answers are not scattered at random. They are the weak half of a clear pair: answers that are vague or wrong, where a judge has to decide how much to forgive. The wording of the rubric is that decision.
Which is exactly why a reword looks like a regression. It moves the cases you added to the set to catch problems.
Two numbers and a fraction. The spread is three times what a dashboard shows, and the ranking survives it anyway.