paper web signal

MRPO Cuts Medical-VQA Cascading Failures 64%→13%, Lifts 8B Model Past 34B Rival

Summary

Shows that 64% of medical multimodal reasoning errors begin with a flawed early reasoning step that cascades to a wrong answer — and a targeted RL fix reduces that to 13%, letting an 8B model beat a specialized 34B rival. Challenges the assumption that bigger medical foundation models are the primary lever.