huggingface.co web signal

UIUC's SILSA Compresses 3D Generation Into 384 Slice Latents

TL;DR

  • UIUC's SILSA tokenizes 3D shapes into just 384 sliding-window slice latents, 128 per axis across three canonical axes.
  • The paper reports PSNR of 32.74 vs 30.12 for SparseFlex, with 70% fewer tokens than the next-most compact baseline.
  • Training memory drops 40.4% and inference falls to 0.34 seconds per shape, a reported 58.5% speedup over the strongest baseline.

A new University of Illinois Urbana-Champaign paper proposes squeezing a 3D shape into 384 tokens. That is 128 slice latents along each of the three canonical axes, used as the entire input to a single-stage rectified-flow generator. The comparable SparseFlex pipeline reports 87,453 tokens on average.

The authors call the method SILSA, for sliding-window slice latents. "Instead of generating expensive voxel tokens, SILSA uses a fixed set of overlapping slices along the three canonical axes, where each token summarizes a local depth window to preserve cross-sectional continuity and support single-stage rectified-flow generation," the paper writes. A Slice VAE encodes the slices, a sparse volumetric decoder reconstructs them, and a Volumetric Anchor Lattice coordinates the three axis streams through a shared 3D workspace. Topology is held in place by two extra losses: one that matches persistence diagrams per slice, and one that aligns Betti transitions across neighbors.

The reported deltas against the strongest baseline are PSNR 32.74 vs 30.12, coverage up 5.96 absolute points, and Betti error down 9.2%. Training memory falls from 55.4 GB to 8.7 GB (a 40.4% reduction as the paper measures it), and inference time drops to 0.34 seconds per shape, which the authors frame as 58.5% faster. SILSA runs at 96M parameters against SparseFlex's 213M.

It lands in a busy week for generative-3D and compression research on our radar, alongside the World Observer video-world-model paper from KAIST we covered the same day. The abstract reports gains on "thin structures, repeated components, and long-range connectivity" but publishes no breakdown by object class.