KAIST SQAM Beats OGBench's Hardest Tasks by 18-35 Points
TL;DR
- SQAM replaces per-step vector-Jacobian products with a scalar adjoint equal to the critic's gradient at the final action, scaled by flow time.
- On OGBench's four hardest domains, SQAM exceeds the strongest baseline by 18 to 35 percentage points, including +35 on cube-quadruple.
- On a Rainbow Robotics RB-Y1 bimanual rig, SQAM online hits 25/30, 15/30 and 13/30 across three tasks versus supervised fine-tuning's 15, 7 and 4.
The batch-averaged velocity Jacobian of a pretrained flow policy, it turns out, concentrates on its diagonal. That observation, in a new KAIST paper on Hugging Face titled "Q-Learning with Scalar Adjoint Matching," is what Minsung Yoon, corresponding author Jinwoo Shin and three co-authors from KAIST and RLWRLD use to collapse an expensive piece of flow-policy fine-tuning into a one-line computation.
Flow policies generate actions over many steps, which makes off-policy RL fine-tuning against a learned value function awkward: exact adjoint matching needs a vector-Jacobian product at every flow step. SQAM (Scalar Q-Adjoint Matching) replaces all of that with a scalar adjoint equal to the critic's gradient at the final action, scaled by flow time. "SQAM observes that the batch-averaged velocity Jacobian of pretrained flow policies concentrates on its diagonal," the paper writes, "enabling a closed-form scalar adjoint that scales the value gradient at the final action by flow time, eliminating per-step vector–Jacobian products required by exact adjoint matching."
The scalar adjoint alone underperforms on manipulation. The authors pair it with a value penalty at policy-generated actions, default coefficient 0.3, switched off for cube-double.
On OGBench, "gains concentrate on the four hardest OGBench domains, where success rate exceeds the strongest baseline by 18 to 35 percentage points." Those four, versus TRQAM: antmaze-giant 62 vs 41, humanoidmaze-large 54 vs 36, cube-triple 72 vs 50, cube-quadruple 54 vs 19.
The authors then port the method to a real robot: a Rainbow Robotics RB-Y1 bimanual manipulator with two 20-DoF Wuji Hand 2 dexterous hands and a Stereolabs ZED Mini stereo camera, running the RLDX-1 vision-language-action model with 16-step action chunking. Across 30 episodes per task, SQAM's offline→online variant lands 25/30 flipping a plastic bag, 15/30 placing a straw in a cup, and 13/30 putting fruit in a pot and closing the lid. Supervised fine-tuning gets 15, 7 and 4.
Code is at github.com/yonghdong/sqam. It is KAIST's second research alert we've tracked this month, after World Observer on October 2.
Originally reported by huggingface.co
Read the original article →Original headline: KAIST SQAM Replaces Per-Step Jacobians in Flow-Policy RL With a Closed-Form Scalar Adjoint, Hits 75% on OGBench