Decision models lack reliability in long chains
New benchmark JEVal reveals that while general decision models are efficient for single-step judgments, their reliability drops in multi-step interactions due to error accumulation, and they exhibit systematic bias in social simulations.
ImportanceLocalEvidenceE2 unreplicatedWrite-upQuick
General decision models like Jev weaken when tasks require specialist knowledge or faithful uncertainty estimation, often overstating outcome probabilities.
In long-horizon, multi-step tests such as τ-bench, faster local decisions reduce median episode time but lower overall task success as errors accumulate over long trajectories.
Large-scale social simulations show these models approach generative LLMs in individual response prediction at lower cost, yet remain weaker in user profiling and exhibit larger aggregate estimation errors and systematic bias.