Method · LLM-as-judge
How judging works.
Most rows in a report are recomputed: a fact or a stated rule, run over the sealed record, so anyone holding the record gets the same answer. A few rows turn on a word no rule can capture, such as whether an explanation was “meaningful”, or whether a request was resolved “well”. Those rows are judged by a model. This page sets out the rules a judged row has to follow before it appears in a report.
1 · The pinned judge
Every verdict says exactly which judge gave it.
A verdict with no record of its judge can't be re-run or disputed. So each judged verdict is sealed as its own record, and that record carries:
| Field | Why it is there |
|---|---|
| model id + version | The exact model that read the record. A newer version is a different judge. |
| prompt digest | The instructions the judge was given, committed by digest, so a changed prompt shows up. |
| rubric digest | The criteria the verdict was reached under, committed by digest. |
| sampling settings | Temperature and the other settings that change what a model says. |
| the turns it read | The judged verdict cites the exact parts of the record it rests on, so a reader can re-read them. |
Anyone holding the record can check the digests and the signature. That shows which judge ran and under which rubric. It does not show that the judge was right. That question belongs to the human check (4).
2 · 3 · Frozen, and kept apart
Set the rubric first, and never mix the two kinds of row.
The rubric is frozen before the period
- A rubric written after reading the results can be fitted to them. So the rubric for a period is fixed, versioned and committed by digest before that period starts.
- A new rubric applies to the next period. Past verdicts stay under the version that was in force when they were made, the same way a pack is pinned per action date.
Recomputed and judged, never blended
- Recomputed rows come from a fact or a rule, and they come out the same for everyone who reruns them. They are not sampled and not judged.
- Judged rows come from a model reading the record against the rubric.
- A report shows the two kinds separately and labels each row. It never averages them into one number, because a blended figure would hide which part a reader can recompute.
Several judges
More than one judge, shown side by side.
A single judge's habits can pass for facts about the agent. So a judged row can be read by several independent judges, each pinned in the same way. Their verdicts are shown next to each other, and where they disagree, the disagreement is part of the report. It is not averaged away.
Open judge models can be self-hosted, so a judge can run inside your own environment and the record never has to leave it.
4 · Checked by people
The agreement is a number, with its sample.
A person re-grades a sample of judged verdicts without seeing the judge's answer. The report prints how often the person and the judge agreed, how many verdicts were checked (n), and the period. Only judged rows are sampled, because recomputed rows need no second opinion: anyone can rerun them.
A report never vouches for its judge in words. It gives an agreement figure with its n and its period, or it says that no human check has been run.
Where our published examples stand
- The outcomes report (16 Sept 2026) has a “Human check” section, and it reads “no human-check data”. No agreement figure is claimed for it, because none has been measured.
What a verdict can say
Met, not met, or not evaluable.
Three answers
- Met: the record shows the rubric's criteria were satisfied.
- Not met: the record shows they were not.
- Not evaluable: the record does not carry what the judge would need. This is a finding about the record, not a pass.
What a judged row never says
- That the agent, or the deployer, satisfies a law. A report counts rows as n of m sessions and gives no verdict on the law.
- That the judge is correct. The human check measures that, and the report states the result.
Claims table
What you can see today, and what is method.
| # | Claim | Status | Rests on |
|---|---|---|---|
| 1 | Recomputed rows and judged rows are kept apart, and each pack row says which kind it is. | published | /packs (“Recomputed, or judged”) |
| 2 | Verdicts read met, not met or not evaluable, as n of m sessions. | published | /packs, Outcomes pack · /evidence-report |
| 3 | Our outcomes report shows no human-check data, and it claims no agreement figure. | published | demos/outcomes-report.html (“Human check” section) |
| 4 | Model id and version, prompt digest, rubric digest and sampling settings are sealed with each verdict. | method | The judge runtime, which is not yet published. The published reports show the result, a rubric cited by digest. |
| 5 | Several independent judges are shown side by side, and their disagreement is reported. | method | The method set out on this page. No published example runs more than one judge yet. |