Multi-judge voting tops three languages in hallucination detection
Multi-judge character-level voting fusion ranks first in three of four languages in SHROOM-Visions 2026 hallucination span detection, self-reported by the team.
ImportanceLocalEvidenceE2 unreplicated
Using multiple fine-tuned vision-language models as independent judges and fusing hallucination span predictions by character-level majority voting ranks first in three of four languages in SHROOM-Visions 2026 hallucination span detection, and places in the top three across all languages and metrics.
Previously, single models predicted hallucination spans directly and disagreement between models went unused; the authors state that model disagreement tracks human annotation disagreement.
The results are team self-reported, from Toqeer Ehsan et al.'s arXiv preprint (submitted September 15, updated to v2 on the 16th, accepted to UncertaiNLP 2026@EMNLP), with no task-organizer leaderboard corroboration yet.
Original: arXiv:2609.17327 ↗