arxiv.org web signal

World models verify vision-based control with 130x fewer params

TL;DR

  • Stochastic world models with physically grounded latents act as perception surrogates for formally verifying vision-based neural controllers.
  • The surrogates reproduce held-out frames more faithfully than GAN surrogates carrying up to 130 times as many parameters.
  • On an emergency-braking benchmark, the pipeline resolves the full state space where prior GAN work left 38% unresolved.

Stochastic world models trained with physically grounded latents reproduce held-out driving frames "more faithfully than GAN surrogates with up to 130 times as many parameters," according to a new preprint from I. Samuel Akinwande, Mykel J. Kochenderfer and Clark Barrett. The surrogates are built to plug into standard neural-network verifiers so a controller's behavior under vision input can be analyzed formally, not only tested.

On an emergency-braking benchmark, the authors report resolving the entire state space, improving on earlier GAN-surrogate work that "left unresolved" 38%. On the RGB version of the same benchmark, where "no verification results have previously been reported," their pipeline resolves more than 80% of the state space by combining falsification, adaptive refinement, symbolic and backward analyses.

The abstract puts the premise plainly: "Verifying a vision-based neural feedback system requires a model of the observations its controller acts upon." Two researchers we track shared the preprint shortly after it went up.

Shared on Bluesky by 2 AI experts