Ask an AI to pick the better of two answers. Then ask again with the two answers swapped.
Sometimes it picks a different winner, and nothing about the answers changed.
That's the catch inside LLM-as-a-judge: using one large language model (the kind of AI behind ChatGPT) to grade the output of another, instead of paying a person to read it all.
The setup is simple. You give the judge the original question, the answer (or two competing answers, A and B), and an instruction like "which response is more helpful and factually correct, and why?" It sends back a score or a verdict.
Why would that work at all? Grading an answer is closer to reading comprehension than to writing a new one from scratch, and reading is something these models are already decent at. The 2023 paper that made the technique mainstream (Zheng et al.) found GPT-4's verdicts matched human preferences over 80% of the time, about as often as two humans agree with each other.
That matters if you build AI agents, programs where a model takes actions on its own, like calling tools. Once one runs thousands of times a day, nobody can read every step it took. A judge model becomes the automated check on whether it picked the right tool and gave a good final answer.
The same paper also found position bias: judges often favor whichever answer sits first, regardless of quality. Even GPT-4 kept the same winner after a swap in only 65% of their test cases.
So never trust a single pass. Run it twice with the order flipped, and if the verdict flips too, you were measuring reading order, not quality. Call it a tie.
Quick check before you scroll: Two answers are equally correct, but one is three paragraphs longer. Why might an LLM judge still score it higher, and what's the usual fix?
Full breakdown + the answer: frankduah.me/learnings/2026-10-07-llm-as-a-judge-using-a-model-to-grade-a-model
New here? I post a bite-size AI / ML concept like this every day. Follow me for the daily drop, and it compounds fast. Why I do it: https://lnkd.in/gK8knHDH
#LLMAsAJudge #LLMEvaluation #AIEvaluation #AI #LLM #AIAgents #MachineLearning
The answer
This is verbosity bias: judges tend to associate length with thoroughness even when it's padding. The fix is to instruct the judge explicitly to penalize unnecessary length, or to normalize scores against response length.