World models verify vision-based control with 130x fewer params
TL;DR
- Stochastic world models with physically grounded latents act as perception surrogates for formally verifying vision-based neural controllers.
- The surrogates reproduce held-out frames more faithfully than GAN surrogates carrying up to 130 times as many parameters.
- On an emergency-braking benchmark, the pipeline resolves the full state space where prior GAN work left 38% unresolved.
Stochastic world models trained with physically grounded latents reproduce held-out driving frames "more faithfully than GAN surrogates with up to 130 times as many parameters," according to a new preprint from I. Samuel Akinwande, Mykel J. Kochenderfer and Clark Barrett. The surrogates are built to plug into standard neural-network verifiers so a controller's behavior under vision input can be analyzed formally, not only tested.
On an emergency-braking benchmark, the authors report resolving the entire state space, improving on earlier GAN-surrogate work that "left unresolved" 38%. On the RGB version of the same benchmark, where "no verification results have previously been reported," their pipeline resolves more than 80% of the state space by combining falsification, adaptive refinement, symbolic and backward analyses.
The abstract puts the premise plainly: "Verifying a vision-based neural feedback system requires a model of the observations its controller acts upon." Two researchers we track shared the preprint shortly after it went up.
Shared on Bluesky by 2 AI experts
Originally reported by arxiv.org
Read the original article →Original headline: Stochastic World Models for Verifying Vision-Based Neural Feedback Systems