Simulated study claims cognitively diverse AI juries resist attacks better
Author-run simulations show persona-diverse juries had 10% lower error rates than provider-diverse ones; small sample, unreplicated.
Original event 2026-09-25
A simulation study claims AI juries diversified by cognitive reasoning style are more robust to adversarial judge hacking than juries diversified by model vendor, with a 10% lower error rate.
Author Anya Habana published the post on LessWrong as part of BlueDot's Technical AI Safety Project Sprint. The experiment used 120 self-generated false-belief scenarios, comparing juries of one model prompted into different reasoning personas against juries of models from OpenAI, Mistral and Google.
The post also reports that prompting alone leaked existing model capabilities, while LoRA fine-tuning gave a cognitively diverse jury over 4% higher accuracy than single judge models. All figures are the author's own tests on Gemini-generated scenarios; the sample is small and the findings await independent replication.