NAMVIS Swaps Diffusion for Autoregression, 3x Faster Views
TL;DR
- NAMVIS replaces diffusion with geometry-conditioned next-scale autoregression and reports running over 3 times faster than evaluated diffusion baselines.
- The model samples discrete visual tokens in coarse-to-fine scale steps, with all tokens within a scale and across target views decoded in parallel.
- Across Objaverse, GSO, and OmniObject3D, the authors claim wins over diffusion baselines on PSNR, SSIM, and LPIPS under the same evaluation setting.
A new arXiv preprint argues that sparse-view 3D synthesis does not need diffusion at all. The authors introduce NAMVIS, which they describe as 'a diffusion-free framework that reformulates multi-view image synthesis as geometry-conditioned next-scale autoregression.' The pitch is throughput: 'NAMVIS outperforms diffusion-based baselines in PSNR, SSIM, and LPIPS, while running over 3 times faster than the evaluated diffusion baselines under the same evaluation setting.'
The mechanism swaps iterative denoising for token prediction. NAMVIS 'predicts discrete visual tokens through a small number of coarse-to-fine scale steps, while sampling all tokens within each scale and across target views in parallel.' A component the paper calls Multi-scale Projective Pose Encoding 'injects source and target camera transformations into both target-view self-attention and source-to-target cross-attention at every resolution,' which is how the model keeps views geometrically aligned without the repeated denoising passes diffusion depends on.
The benchmarks are Objaverse, GSO, and OmniObject3D. The abstract names those three evaluation sets but does not identify which diffusion models the comparison runs against, and prints no absolute numbers for the PSNR, SSIM and LPIPS wins or for the per-view latency behind the 3x figure. The authors frame the finding as a direction rather than a settled benchmark: 'geometry-conditioned next-scale autoregression is a promising and efficient alternative to diffusion for sparse-view multi-view synthesis.'
Originally reported by paper
Read the original article →Original headline: NAMVIS Replaces Diffusion for Multi-View Synthesis, Runs 3x Faster at NeurIPS 2026 Quality