vLLM v0.28.0 Ships Decode Context Parallel, Bigger Batches and Model Runner V2 Maturation
Summary
vLLM v0.28.0 landed with 584 commits from 270 contributors: Decode Context Parallel and fused kernels for Kimi-K3, end-to-end sparse MLA for DeepSeek V4, DFlash2 speculative decoding with confidence-scheduled verification, tiered KV cache offloading to disk, and Model Runner V2 maturation. Default max_num_batched_tokens doubles to 16384 and Blackwell CUDA graph capture rises to 1024.
Originally reported by github.com
Read the original article →Original headline: vLLM v0.28.0 Ships Decode Context Parallel, Bigger Batches and Model Runner V2 Maturation