CMU's ModAR Ships First World-Action Model That Autoregressively Denoises Depth, Point Tracks and DINO Features
Summary
CMU researchers Adam Hung, Bardienus Duisterhof, Deva Ramanan and Jeffrey Ichnowski unveil ModAR, described as the first world-action model that sequentially denoises depth maps, point tracks and DINO features before predicting actions rather than focusing on RGB. The team reports 75% success on real-world bimanual robotic tasks while using about 20x fewer training FLOPs than a pretrained Flex-pi baseline, and finds that adding RGB prediction does not consistently improve performance.
Originally reported by huggingface.co
Read the original article →Original headline: CMU's ModAR Ships First World-Action Model That Autoregressively Denoises Depth, Point Tracks and DINO Features