DSReg Recovers Individual JEPA Latents Without Decoder or Labels
TL;DR
- DSReg claims to recover individual world latents from JEPA up to signed permutation with no reconstruction, no decoder, and no labels.
- Its Structural Diversity assumption is argued to be strictly weaker than Structural Sparsity, Non-Inclusion, and Structural Variability from prior nonlinear-ICA work.
- On learned convolutional encoders DSReg matches a supervised Procrustes oracle; on nested-footprint synthetics it reports roughly 0.95 to 0.99 MCC.
Individual world latents can be provably recovered from JEPA representations with no reconstruction, no decoder, and no labels. That is the central claim of DSReg, a paper by Yujia Zheng (University of Illinois Urbana-Champaign), David Klindt (Cold Spring Harbor Laboratory), Randall Balestriero (Brown University), and Bernhard Schölkopf (Max Planck Institute for Intelligent Systems, ELLIS Institute Tübingen).
The method rests on a condition the authors call Structural Diversity: "different latents leave distinct dependency footprints on observations, just as no two snowflakes are alike." Building on the linear identifiability of LeJEPA, DSReg fits an orthogonal rotation on top of a frozen, whitened representation via a dependency-sparsity regularizer, and under that condition it recovers each latent up to signed permutation. "It applies post hoc to any linearly identified representation, reusing trained checkpoints at no loss over joint training," the paper states, and claims to establish "the first fully identifiable JEPA that recovers every world latent."
The reported numbers come from synthetic regimes and small learned encoders. On nested-footprint benchmarks spanning latent dimension N=4 to N=32, DSReg reports recovery of roughly 0.95 to 0.99 MCC where LeJEPA alone remains mixed. On convolutional encoders trained from scratch it "matches supervised Procrustes oracle," while PCA, Varimax and FastICA "leave latents nearly as mixed as unrotated baseline." On Gaussian 3DShapes it rises "from mixed to near-ceiling recovery at matched dense R²," and on quarter-orientation dSprites it beats β-VAE and β-TCVAE. Proposition 1 argues the Structural Diversity condition is "strictly weaker than all structural conditions of prior identifiable latent variable models," including Non-Inclusion and Structural Variability.
Reported scaling reaches d=8192 on a single 48 GB GPU with up to 10^6 samples. The abstract-level evidence does not test the regularizer on transformer-scale JEPA stacks.
Originally reported by huggingface.co
Read the original article →Original headline: DSReg Provably Recovers Individual World Latents Without Reconstruction, Beats Prior Structural Conditions