KAIST SQAM Replaces Per-Step Jacobians in Flow-Policy RL With a Closed-Form Scalar Adjoint, Hits 75% on OGBench
Summary
KAIST and RLWRLD propose Scalar Q-Adjoint Matching (SQAM), observing that batch-averaged velocity Jacobians in pretrained flow policies concentrate on the diagonal — so a scalar adjoint scaled by flow time replaces expensive per-step vector-Jacobian products. On OGBench's 50 tasks it hits 75% success versus 68% for TRQAM and delivers +18-35 points on the hardest cube and humanoid-maze domains, with real-robot bimanual gains on three manipulation tasks.
Originally reported by huggingface.co
Read the original article →Original headline: KAIST SQAM Replaces Per-Step Jacobians in Flow-Policy RL With a Closed-Form Scalar Adjoint, Hits 75% on OGBench