paper web signal

LiFT loops shared DiT core, beats DiT-XL/2 FID by 3.34 points

TL;DR

  • LiFT-L/2 beats DiT-XL/2 on ImageNet 256x256 by 3.34 FID points while running roughly 60% fewer parameters.
  • The looped model also reports 32% fewer training FLOPs and 52% fewer inference FLOPs than the dense baseline.
  • A continuous depth index lets a trained LiFT model loop beyond its training depth at inference with no retraining or early exits.

LiFT-L/2, a looped diffusion transformer, posts an FID 3.34 points below the dense DiT-XL/2 baseline on ImageNet at 256x256 while using roughly 60% fewer parameters, 32% fewer training FLOPs and 52% fewer inference FLOPs, according to a preprint on arXiv.

The design reuses a single DiT core across many recurrent steps rather than stacking distinct layers. 'Rather than asking every recurrent step for the final prediction,' the authors write, 'LiFT trains each step with a single regression target: a point on a straight path from the model's initial estimate to the flow-matching target.'

Because each step is indexed by a continuous depth coordinate, a trained model 'can loop far beyond its training depth with no retraining, early exits, or other modifications,' the paper says, and the authors report that in those longer rollouts 'inference computation can grow without adding parameters.' The abstract reports results on one benchmark, ImageNet at 256x256, with Mohammad Mahdi Derakhshani listed first on the author line.