Cognitively diverse AI juries cut error rates by 10%
Author-run tests show reasoning-style-diverse juries had 10% lower error rates than vendor-diverse ones, with LoRA fine-tuning adding over 4% accuracy; small sample, unreplicated.
중요도국소적증거E2 미복제작성 방식간략
AI juries diversified by cognitive reasoning style are more robust to adversarial judge hacking than juries diversified by model vendor, with a 10% lower error rate — a result the author, Anya Habana, obtained in her own tests.
Jury diversity has typically meant grouping models by vendor, with no direct comparison of reasoning-style diversity; this exploratory work was done under BlueDot's Technical AI Safety Project Sprint and published on LessWrong.
The experiment used 120 self-generated false-belief scenarios (generated by Gemini Flash Lite 3.5), comparing juries of one model prompted into different reasoning personas against juries of models from OpenAI, Mistral and Google, yielding the 10% gap. The post also reports that prompting alone leaked existing model capabilities, while LoRA fine-tuning gave a cognitively diverse jury over 4% higher accuracy than single judge models. All figures are the author's own tests; the sample is small and the findings await independent replication.