New architectures beating old models may not be the architecture's doing
On benchmark leaderboards, new Transformer-based image models have overtaken old-style convolutional networks that look at images through small local windows, and people credit the architecture.
On benchmark leaderboards, new Transformer-based image models have overtaken old-style convolutional networks that look at images through small local windows, and people credit the architecture. The paper modernizes the old convolutional network in the same way as its rivals, keeping only the local-view-of-image part. On the most commonly used image test set it reaches 87.8% accuracy, and it also surpasses the rivals in detection and segmentation.
Next time you see a claim that "a new architecture crushes old methods," first ask whether the training setups on both sides were aligned: amount of data, training duration, augmentation methods. Any gap from unaligned setups should first be chalked up to training investment, before discussing which architecture is better.
If the task is not this kind of standard image benchmark — for example, very high-resolution inputs — don't use it to judge architectures. The paper also doesn't prove that the unmodernized old network was competitive all along; what's competitive is the modernized one.
A ConvNet for the 2020s (2022) | Next review 2027-09-20