Researchers propose monitoring closed models via open-weight internals
First data for probing closed models without weight access, but all results are self-run and real audit use remains unverified.
It may be possible to monitor deception and other misalignment in closed models through the internals of an open stand-in, without ever touching the closed model's weights.
A team led by conti published model hermeneutics on LessWrong on September 30: the closed model under study is the author, an open-weight reader reads its outputs, and linear probes are trained inside the reader. In the authors' own tests, a 27B reader probing a 397B author showed an average AUROC gap of only 0.004.
The team reports probes detecting reward hacking, sycophancy and deception in outputs from three frontier models including Claude Opus 4.6, Gemini 3.1 Pro Preview and GPT-5.4, beating direct reader interrogation in roughly half the settings. Distillation lifted the reader's deception-probing AUROC from 0.69 to 0.90.
This is roughly 1.5 months of early work; closed-model internals have no ground truth, so faithfulness is measured only against indirect open-weight baselines, and real audit use awaits independent verification.