Convolution losing to attention may just mean the training recipe lost
Seeing vision Transformers beat ResNet on leaderboards, commentary often attributes it to attention being inherently stronger.
Long-term reading
Seeing vision Transformers beat ResNet on leaderboards, commentary often attributes it to attention being inherently stronger. This work modernizes an old convolutional network item by item according to its rival's training recipe, and the resulting ConvNeXt achieves 87.8% on ImageNet and surpasses Swin on detection and segmentation.
Next time you encounter claims that "a certain architecture has been made obsolete," first check whether both sides of the comparison used the same generation of training recipe. Architectural merit and training investment should be accounted for separately — falling behind may just mean it was never modernized.
If the task isn't a benchmark like ImageNet, COCO, or ADE20K, or the input resolution is extremely high, don't use it to judge whether convolution or Transformers are stronger. It also only speaks for the modernized convolutional network — don't use it to judge the original ResNet.
《A ConvNet for the 2020s》(2022) | Next review 2027-09-20