arxiv.org web signal

Latent-Foresight jointly trains VFM tokenizer and dynamics model

TL;DR

  • Latent-Foresight jointly trains a latent tokenizer and a flow-based generative dynamics model, replacing two-stage pipelines that use PCA or a frozen autoencoder.
  • The authors claim consistently better results than two-stage baselines across multiple future scene understanding tasks and prediction horizons, but publish no numbers in the abstract.
  • Code and model weights are released on GitHub under a CC-BY 4.0 license.

A new preprint from Efstathios Karypidis, Spyros Gidaris, and Nikos Komodakis proposes training the latent tokenizer and the dynamics model together, rather than compressing Vision Foundation Model features with PCA or a frozen autoencoder and bolting a predictor on top.

The method, called Latent-Foresight, is described in the arxiv abstract as 'an end-to-end framework that jointly learns a latent tokenizer and a flow-based generative dynamics model, explicitly shaping the representation to support temporal predictability.' The authors say several design choices 'prevent latent collapse and align reconstruction with generative objectives' during joint optimization.

The headline claim is that the shared objective 'consistently outperforms two-stage baselines across multiple future scene understanding tasks and prediction horizons.' The abstract names no benchmarks and publishes no numbers. It does not list the prediction horizons tested. Code and weights are released under CC-BY 4.0, and the paper landed the same day as another latent-compression result we tracked for 3D generation.