Five frontier models checked the same claims and disagreed on most
A self-run study finds large disagreement and overconfidence across models fact-checking real claims; AI verdicts are not interchangeable.
Original event 2026-09-24
A self-run study had five frontier LLMs fact-check the same 997 real-world claims, and they disagreed on more than 60% of them.
The study was posted on LessWrong by the team behind the fact-checking platform Lenz.io, using claims users actually submitted between May and July 2026 rather than public benchmark questions. The five models agreed unanimously on only 37% of claims, and on 23% the most distant verdicts differed by at least two categories on a five-point scale.
The models were broadly overconfident: 76% of answers reported confidence of 9 or 10 out of 10, yet disagreement still ran at 63%. The team states it did no human labelling, so it could measure only disagreement, not which model is more accurate.
The usable takeaway: fact-check verdicts from different models cannot be swapped for one another, and one model's high confidence does not mean the others agree.