Interpretability tools get a new exam, with hallucination rates now scored
Researchers built and open-sourced WorkspaceBench to test activation-reading tools; it is a self-built benchmark, and independent adoption remains to be seen.
ImportanceLocalEvidenceE2 unreplicatedWrite-upQuick
Interpretability researchers have released and open-sourced WorkspaceBench, a benchmark testing whether activation-to-text tools can actually read a model's global workspace.
The benchmark has 3,356 questions across 27 eval families covering safety, logical reasoning and multihop computation, plus a dedicated hallucination eval, because J-lens is reliable but single-token while NLAs are expressive yet prone to confabulation.
It was built for Qwen-3.6-27B, and the authors admit it does not fully rule out shortcuts that infer intermediates from the prompt. This is a self-built, self-assessed benchmark; whether the field adopts it remains an open question.