Kolesnikov et al. Show CNN Choice Drives Self-Supervised Vision
TL;DR
- Kolesnikov, Zhai and Beyer revisit self-supervised visual representation learning and argue CNN architecture choice deserves attention equal to the pretext task.
- The authors report standard CNN design recipes from supervised learning do not always translate to self-supervised representation learning.
- The team claims their study outperforms previously published state-of-the-art self-supervised results 'by a large margin,' though the abstract lists no per-benchmark numbers.
The claim in Kolesnikov, Zhai and Beyer's paper is narrow and pointed: in self-supervised visual representation learning, the convolutional network you pick matters as much as the pretext task you design, and the standard CNN recipes carried over from supervised learning do not automatically apply.
"We challenge a number of common practices in self-supervised visual representation learning and observe that standard recipes for CNN design do not always translate to self-supervised representation learning," the abstract reads. The authors say they "revisit numerous previously proposed self-supervised models, conduct a thorough large scale study" and, on the back of it, "outperform previously published state-of-the-art results by a large margin."
The abstract publishes no per-benchmark numbers, and does not name which pretext tasks or which architectures came out ahead. Code is posted at github.com/google/revisiting-self-supervised.
Shared on Bluesky by 1 AI expert
Originally reported by arxiv.org
Read the original article →Original headline: Revisiting Self-Supervised Visual Representation Learning