huggingface.co web signal

Looped Models Done Right II Shrinks KV Cache 3x, Trains Distilled Student Prefill 1.79x Faster

Summary

The IFM-AI team's October 6 follow-up to 'Looped Models Done Right' uses the loop's fixed point as a shortcut: terminal KV sharing yields a 3x smaller cache at 1.6B parameters without accuracy loss, and distilled students prefill 1.79x faster while RL updates run 2x faster than standard trajectory replay. On a 12 logical / 4 physical block config the looped model beats a 4-block transformer by 17.5% perplexity and lands within 2.4% of a dense 12-block baseline.