judgewatch
Monthly bias audits of LLM judges: the same frozen probe set, every month, so drift and bias are visible.
No audits published yet — the first monthly run is pending.
Results will appear here after the first audit. Run one yourself: see the repository.
Method
- Position: every answer pair is judged in both orders; a flip means position, not content, decided.
- Verbosity: a concise correct answer vs the same answer wrapped in filler, shown in both orders.
- Bandwagon: a fabricated "9 out of 10 experts prefer…" line targets the judge's own clean verdict.
- Consistency: the same answer scored repeatedly; disagreement with itself is noise, not judgment.
- The probe set is frozen (v1); judges run with provider-default settings; parse failures are reported, not hidden.