Study says small models notice being tested, but most don't change behavior
Verbalized eval awareness is largely not causal for refusal in open models, so fears of distorted safety evals may be overstated.
A first-party study finds that open reasoning models often say "I might be being tested" in their chain of thought, yet most do not change their behavior because of it.
Sixteen open reasoning models were tested with the Fortress safety prompt dataset. Rates of verbalized eval awareness ranged from very rare to about a third of rollouts and were largely unrelated to model size; within the same question, rollouts containing such statements refused more often, in 14 of 16 models.
The causal tests point the other way: inserting "this might be a test" raised refusal about as much as any sentence about being manipulated, and after removing the model's own eval-awareness sentences, only Nemotron 3 Super and Qwen3 32B changed behavior. The authors note this does not strictly prove causality, and the findings may not extend to frontier closed models.
Sources:https://www.lesswrong.com/posts/Pd5fwZ9Ct5p6hAkrA/evaluation-awareness-in-small-ish-models