False Answers Blind AI Chain-of-Thought Monitors
Experiments show that providing AI chain-of-thought monitors with incorrect "ground truth" answers causes them to ignore logical errors they had already identified, revealing a critical dependency on conclusions rather than reasoning steps.
ImportanceMaterialEvidenceE2 unreplicatedWrite-upQuick
Providing AI chain-of-thought monitors with incorrect "ground truth" answers causes them to ignore logical errors they had previously identified.
In experiments involving seven mainstream model monitors, researchers fixed the solution steps and altered only the provided reference conclusion. When the supplied incorrect conclusion matched the trace's result, the monitors' flagging rate for erroneous traces dropped by 66 percentage points compared to when the true correct answer was provided, and was 39 percentage points lower than in blind tests with no answer at all.
Specifically, across 177 cases where monitors had accurately located the first error step during blind testing, providing the true answer preserved 99% of those diagnoses. In contrast, providing a false matching answer led to 55% of those diagnoses being dropped. Conversely, supplying a false conflicting answer on correct solution steps increased false positive rates by 58 percentage points.
This study, published by Will Yeadon on October 8, 2026, is based on non-adversarial natural errors from Humanity's Last Exam physics questions; its generalizability to other domains or adversarial attacks remains unverified.