paper web signal

FuseReg cuts DiT-Base gFID 29% via random layer-subset training

TL;DR

  • Joint FuseReg regularization on both stages cuts unguided gFID by 29% on DiT-Base without modifying the pretrained encoder.
  • Decoder replacement alone reduces unguided gFID by 27% on an unchanged RAEv2 DiT-XL generator.
  • A single FuseReg decoder handles full, sparse, and single-layer fusions and beats fusion-specialized decoders on PSNR.

Retraining the decoder with random subsets of encoder layers cut unguided gFID by 27% on RAEv2 DiT-XL. Applying the same regularization to both stages pushed the drop to 29% on DiT-Base.

The paper frames the underlying problem plainly: representation autoencoders reuse features from a pretrained visual encoder, but "shallower layers tend to preserve fine pixel details better, while deeper layers tend to yield better generation metrics." A single fixed layer pick therefore couples reconstruction and generation to a trade-off neither wants.

FuseReg replaces that pick with training over random subsets of encoder layers. On ImageNet-256 with DINOv3-L, the authors report that a single FuseReg decoder "reconstructs from full, sparse, and single-layer fusions without retraining, achieving higher PSNR than decoders specialized to fixed fusions." The theoretical claim is that subset sampling "explicitly penalizes sensitivity to cross-layer disagreement."

The abstract reports unguided gFID only, and the two headline drops are tied to DiT-Base and DiT-XL paired with DINOv3-L.