Seeing What a Neural Network Actually Recognizes: Activation Atlases Draw a Map of Concepts
Some image recognition models learn a pile of concepts internally, but no one can say exactly what they recognize.
Some image recognition models learn a pile of concepts internally, but no one can say exactly what they recognize. Activation Atlases average the model's internal responses to millions of images and draw them as a concept map, revealing the model's high-level misunderstandings: for example, a baseball photo led the model to recognize a "gray whale" as a "great white shark."
Today, when debugging why an image model misrecognizes something, this kind of visualization is still a standard tool; most later interpretability methods are patches built on top of it. The trigger scenario is a model misjudging something and you wanting to see what it internally took the image to be.
If the model isn't an image classifier, or the images are far from the distribution of common test sets, don't use it to judge model behavior. It draws a low-dimensional projection whose paths don't necessarily correspond to real structure — treat it as a clue, not quantitative proof.
《Activation Atlases》(2019) | Next review 2027-09-20