Opening up neural networks to read out their principles gets harder as models grow larger
Neural networks are the kind of programs behind chatbots and image recognition, internally made of massive numbers of mutually weighted digits, with no one having written the rules line by line.
ImportanceMaterialEvidenceE3 inspectableWrite-upQuick
Neural networks are the kind of programs behind chatbots and image recognition, internally made of massive numbers of mutually weighted digits, with no one having written the rules line by line. In a long interview, Chris Olah discussed the direction he is betting on: mechanistic interpretability—opening up the network, reading out the steps it actually performs, and giving an account of each step that can be tested and potentially overturned.
Today, when AI companies talk about safety, they often say they want to "open up the model and look inside," following exactly this program. The objection he himself worries about most in the interview is that as model scale grows, this kind of disassembly may not keep up, and this remains unresolved to this day. When you see claims like "we understand our own models," first ask: how large a model did they actually open up.
The words in the interview are expectations about a research direction, not verified conclusions. If someone uses it to assert that some large model has already been understood, or can never be understood, don't use it to judge—that goes beyond what the interview can support.
"Chris Olah on what is going on inside neural networks" (2021) | Next review 2027-09-20