Five frontier models fact-checked the same claims and mostly disagreed
A self-run test by the Lenz.io team found five models disagreed on over 60% of the same claims while reporting high confidence, so verdicts are not interchangeable.
ImportanceLocalPreuvesE2 non réplicableTraitementRapide
Five frontier LLMs fact-checking the same 997 real-world claims disagreed on more than 60% of them — meaning AI fact-check verdicts from different models cannot be swapped for one another.
Previously, a reader relying on a single model's verdict had no way to know whether other models would agree, and the models' self-reported high confidence made the conclusions seem more reliable than they are.
The self-run test was carried out by the team behind the fact-checking platform Lenz.io and posted on LessWrong, using claims users actually submitted between May and July 2026 rather than public benchmark questions. The five models agreed unanimously on only 37% of claims, and on 23% the most distant verdicts differed by at least two categories. The models were broadly overconfident: 76% of answers reported confidence of 9 or 10 out of 10, yet disagreement still ran at 63%. The team states it did no human labelling, so it could measure only disagreement, not which model is more accurate.
The results have not been rechecked against human labels or reproduced by any third party; the sample was user-submitted claims from May to July 2026, posted on LessWrong.