The notary
Commissioned, four year termEvery obligation is written on the form. Break one and the bond pays out.
Sealed
The station at the end of the loop
A dispute arrives carrying everything you need to judge it. Rule on it, then read the form the architecture keeps on you. The notary has been filling the same one out since long before any of this, and the difference is the entire argument.
Open dispute, 2 days remaining
EMEA retention rate
Decision impact class is set to moderate. It was certified eight months ago and it now feeds headcount planning, so the class understates what happens when the number is wrong.
Raised by
M. Okafor, finance. Four disputes filed, three upheld.
Owner response
Held at moderate. Answered inside the deadline.
What each station owes
One profession has answered this for four hundred years. Rule above to print yours next to theirs.
Every obligation is written on the form. Break one and the bond pays out.
Sealed
The same result, measured
A reviewer can answer almost identically every time you ask, and answer with a systematic tilt.
The largest systematic study of LLM-as-judge published so far covers 21 judges from nine providers across 118 runs and roughly 541,000 individual judgments. Judge rankings shift by as much as 14 positions depending on which benchmark ranks them, and two production-deployed judges recorded test-retest reliability above 0.95 alongside position bias above 0.10. A companion paper treats replacing the evaluator as a measurement-validity problem, because the measured result moves even when the material under evaluation has not changed. Consistency is the property behavioral monitoring captures well. It is not the property that told you whether the reviewer was right.
One intervention does work, and it is not more review. When researchers benchmarked the standard quality-control methods against a hard annotation task, only one improved quality. Visible gold questions, which give workers periodic feedback on their own accuracy while they work, for a 7 percent gain over baseline. Concealed scoring did not do it. Letting the checked party see the check did.
Read the full article
Who rates the reviewer, on SubstackThe essay carries the built case, why a second adjudicator was rejected at the ruling, and what notaries settled about this a long time ago. Adjacent, and worth a look if you came for the review surface itself, is the eval-to-label loop.
Get the next one
One essay a week on making data and AI systems trustworthy enough to act on.
Part of: Stage 03 · Grading Itself
Back to the mapA contract that cannot grade itself is decoration. The architecture audits its own trust scores against what actually happened, because a score nobody checks against outcomes is a check engine light that has been on for two years.
Read the essays