01 Aug 5, 2026 · 9 min read
LLM as a Judge: Can a Model Evaluate Test Results, and When Not to Trust It
The LLM-as-judge pattern in testing: a rubric instead of a 1-10 scale, calibration on golden examples, three working modes, and four judge biases you need to know before rollout.