paper web signal

reViT: Single Recurrent Block Matches DeiT III, 70% Fewer Params

TL;DR

  • reViT applies a single transformer block repeatedly to match DeiT III accuracy with roughly 70% fewer stored parameters at comparable inference compute.
  • The feed-forward network is represented as a convex combination of a small shared expert bank, with depth coordinates programming the mixture.
  • An 8-expert version distilled from a DINOv2 teacher recovers nearly all of the teacher's linear-probe accuracy and transfers to segmentation and depth.

A single transformer block, applied over and over, can match the accuracy of a full-depth ViT. That is the claim in a new arXiv paper from Adrian Bulat, Yassine Ouali and Georgios Tzimiropoulos, posted on October 8. Their system, reViT, hits DeiT III accuracy with roughly 70% fewer stored parameters at comparable inference compute.

The trick is in the feed-forward layer. Each depth gets its own mixture, drawn from "a convex combination of a small shared expert bank," with depth coordinates programming which experts activate. The authors report that weight-space merging is "the strongest tested MoE family at a matching one-FFN budget," outperforming token-dispatch and output-mixture alternatives.

An 8-expert version distilled from a DINOv2 teacher recovers nearly all of the teacher's linear-probe accuracy, and the architecture transfers to classification, segmentation and depth prediction. The abstract publishes no per-task benchmark numbers beyond that linear-probe comparison.