FlashPrefill V2 Posts 47x Speedup Over FlashAttention-2 at 128K Context, Merges Into SGLang
Summary
At 47x FP8 speedup over FlashAttention-2 at 128K context with SGLang already integrated, FlashPrefill V2 could immediately shift the economics of long-context LLM serving for any team running on H20 hardware—this is not a research prototype but a production-ready kernel drop-in.
Originally reported by paper
Read the original article →Original headline: FlashPrefill V2 Posts 47x Speedup Over FlashAttention-2 at 128K Context, Merges Into SGLang