paper web signal

FlashPrefill V2 Posts 47x Speedup Over FlashAttention-2 at 128K Context, Merges Into SGLang

Summary

At 47x FP8 speedup over FlashAttention-2 at 128K context with SGLang already integrated, FlashPrefill V2 could immediately shift the economics of long-context LLM serving for any team running on H20 hardware—this is not a research prototype but a production-ready kernel drop-in.