Type Safety Does Not Ensure Logical Correctness
New preprint reveals LLMs often follow option name semantics over definitions, causing high decision flip rates.
ImportanceMaterialEvidenceE2 unreplicatedWrite-upQuick
Even when outputs strictly adhere to type definitions, large language models may deviate from intended logic due to semantic cues in option names.
A new preprint by Yu Sun et al. finds that changing binary decision option names from "0/1" to "yes/no" increases decision-flip rates by up to 70.4 percentage points. Despite a 0% type-error rate, models tend to follow the polarity of the name rather than the bound definition text.
The study tested Jev and two open-weight models on 1200 tasks. Results show that random strings as names perform similarly to neutral controls, indicating that semantically loaded names like "yes/no" introduce instability beyond simple reassignment.
These are author-reported results without independent reproduction yet. The finding impacts prompt engineering and evaluation frameworks, highlighting risks of logical misalignment masked by structural compliance.