arxiv.org web signal

DINO self-supervised ViTs reach 80.1% top-1 on ImageNet

TL;DR

  • DINO's self-supervised ViT features reach 78.3% top-1 on ImageNet as a k-NN classifier with a small ViT, no fine-tuning.
  • Paired with ViT-Base, DINO hits 80.1% top-1 on ImageNet under linear evaluation.
  • The authors flag momentum encoder, multi-crop training, and small patches as the ingredients that make the recipe work.

The 2021 DINO paper from Mathilde Caron and colleagues argues that self-supervised Vision Transformers develop properties their supervised and convolutional counterparts do not. The authors write that the features "contain explicit information about the semantic segmentation of an image, which does not emerge as clearly with supervised ViTs, nor with convnets."

Those same features double as classifiers without fine-tuning. The paper reports they are "excellent k-NN classifiers, reaching 78.3% top-1 on ImageNet with a small ViT." Paired with a ViT-Base, the linear-evaluation number climbs to 80.1%.

The method the authors describe as "a form of self-distillation with no labels." The abstract credits three engineering choices: momentum encoder, multi-crop training, and the use of small patches with ViTs.

The link is back in circulation among the researchers we track; two of them shared it this week.

Shared on Bluesky by 2 AI experts