Rater guidelines
The same definition of the dimension, in writing
Solutions
Where automation cannot see the failure, a human must — trained, calibrated, and measured. Unstructured “eyeballing” is not a programme. Agreement is the score of the scoring.
For high-stakes or subjective dimensions: tone in a regulated reply, clinical or legal nuance, preference between two fluent answers, or the slice an LLM-judge keeps missing.
A rater without a guideline is improvising. We write the guide from the rubric, train on it, and keep a calibration set that is scored again as people drift. New raters do not join a live queue until they match the standard.
Guidelines are short enough to use under time pressure and specific enough to survive a disagreement. If two trained people cannot apply a sentence the same way, the sentence is wrong — not the people.
Inter-annotator agreement is tracked. Disagreements are adjudicated, not averaged away. Preference ranking and pairwise comparison are used when the question is “which is better,” not “is this true.” The method follows the decision.
Adjudication is logged. A later reader should see why the gold label is the gold label. Hidden consensus is how human eval becomes theatre.
Humans are expensive. The design is usually hybrid: automatic scores on the bulk, human review on a sample, on disagreements, and on the critical slice. Cost is an evaluation parameter, not an embarrassment.
The sample is not random courtesy. It is stratified by slice and by the auto-score’s uncertainty. That is how a small panel still sees the cases that would fail a gate.
These artifacts are what make a human score repeatable.
The same definition of the dimension, in writing
Whether people still agree with last month
IAA, adjudication log, rater notes
What the humans see, and what the auto-score covers
Typical shape: Framework + Pilot, or a workstream inside Full Build / retainer, typically designed in weeks; operated on a cadence you keep. Duration is typical, not contractual. Formats and deliverables are written on the evaluation ledger.
Get Started
Write with a system in mind. We will tell you whether it can be evaluated, and what the first ledger would contain.
Request an evaluation