MRPO Cuts Medical-VQA Cascading Failures 64%→13%, Lifts 8B Model Past 34B Rival
Summary
Shows that 64% of medical multimodal reasoning errors begin with a flawed early reasoning step that cascades to a wrong answer — and a targeted RL fix reduces that to 13%, letting an 8B model beat a specialized 34B rival. Challenges the assumption that bigger medical foundation models are the primary lever.
Originally reported by paper
Read the original article →Original headline: MRPO Cuts Medical-VQA Cascading Failures 64%→13%, Lifts 8B Model Past 34B Rival