Skip to content

The station at the end of the loop

Who rates the reviewer?

A dispute arrives carrying everything you need to judge it. Rule on it, then read the form the architecture keeps on you. The notary has been filling the same one out since long before any of this, and the difference is the entire argument.

Open dispute, 2 days remaining

EMEA retention rate

Decision impact class is set to moderate. It was certified eight months ago and it now feeds headcount planning, so the class understates what happens when the number is wrong.

Raised by

M. Okafor, finance. Four disputes filed, three upheld.

Owner response

Held at moderate. Answered inside the deadline.

What each station owes

One profession has answered this for four hundred years. Rule above to print yours next to theirs.

ICertificate of authority
A-1 / 0001

The notary

Commissioned, four year term
Authority fromA commission, revocable
A wrong call costsThe posted bond
Authority lapsesOn the expiry date
Decisions land inThe journal, by entry

Every obligation is written on the form. Break one and the bond pays out.

Sealed

The same result, measured

A reviewer can answer almost identically every time you ask, and answer with a systematic tilt.

The largest systematic study of LLM-as-judge published so far covers 21 judges from nine providers across 118 runs and roughly 541,000 individual judgments. Judge rankings shift by as much as 14 positions depending on which benchmark ranks them, and two production-deployed judges recorded test-retest reliability above 0.95 alongside position bias above 0.10. A companion paper treats replacing the evaluator as a measurement-validity problem, because the measured result moves even when the material under evaluation has not changed. Consistency is the property behavioral monitoring captures well. It is not the property that told you whether the reviewer was right.

One intervention does work, and it is not more review. When researchers benchmarked the standard quality-control methods against a hard annotation task, only one improved quality. Visible gold questions, which give workers periodic feedback on their own accuracy while they work, for a 7 percent gain over baseline. Concealed scoring did not do it. Letting the checked party see the check did.

Read the full article

Who rates the reviewer, on Substack

The essay carries the built case, why a second adjudicator was rejected at the ruling, and what notaries settled about this a long time ago. Adjacent, and worth a look if you came for the review surface itself, is the eval-to-label loop.

Get the next one

One essay a week on making data and AI systems trustworthy enough to act on.

Subscribe on Substack →

Part of: Stage 03 · Grading Itself

Back to the map

A contract that cannot grade itself is decoration. The architecture audits its own trust scores against what actually happened, because a score nobody checks against outcomes is a check engine light that has been on for two years.

Read the essays