Open-weight replication fails to support the global workspace hypothesis
J-lens interventions rarely flip final answers on open-weight models, so the hypothesis's scope needs re-examining.
ImportanceLocalEvidenceE2 unreplicatedWrite-upQuick
A researcher replicated Anthropic's global workspace experiment on open-weight models, and the results did not support the hypothesis.
Anthropic's J-lens paper proposes that a model's intermediate reasoning concepts live in a "global workspace", and that intervening in this space should change the model's answer. Author Mirella Zeisler re-ran the multi-hop reasoning experiment on Qwen3.6-27B and Gemma 3 27B-it — for example, swapping the intermediate step "Mars" for "Neptune" and checking whether the model answers "blue".
The interventions shifted answer probabilities but rarely flipped the output: top-1 flip rates were 6.3% to 11.1% across four conditions, far below the 54% to 70% Anthropic reported on Claude. In three of the four conditions, injecting the counterfactual answer directly beat injecting the counterfactual intermediate, suggesting the intervention may simply steer the model toward its final answer.
The author concludes that J-lens is more useful for reading intermediate variables than for steering model outputs. This is a first-party experiment on a personal blog; after filtering, each condition retained only 27 to 44 prompts, so the samples are small.