Most open models say they're being tested, but don't change behavior
Verbalized eval awareness is largely not causal for refusal in open models, so fears of distorted safety evals may be overstated.
ImportânciaLocalEvidênciaE2 não replicadaTratamentoRápido
Most open reasoning models say "I might be being tested" in their chain of thought, yet do not change their behavior because of it — fears that safety evals are being distorted may be overstated.
The concern had been that models noticing they are being evaluated would alter their answers and skew safety evals; a first-party study tested whether that concern holds.
Sixteen open reasoning models were tested with the Fortress safety prompt dataset. Rates of verbalized eval awareness ranged from very rare to about a third of rollouts and were largely unrelated to model size; within the same question, rollouts containing such statements refused more often, in 14 of 16 models. The causal tests point the other way: inserting "this might be a test" raised refusal about as much as any sentence about being manipulated, and after removing the model's own eval-awareness sentences, only Nemotron 3 Super and Qwen3 32B changed behavior. The authors note this does not strictly prove causality, and the findings may not extend to frontier closed models.