vLLM v0.31.0 released, focused on DeepSeek-V4.1-Flash inference performance
Announced October 6, 2026
What happened
vLLM v0.31.0 includes 717 commits from 307 contributors (96 of them new). The main change is better DeepSeek-V4.1-Flash performance: FlashMLA mega attention with V4.1 NVFP4 compressed KV cache is now the default on SM100. It also adds kernel fusions such as DeepGEMM sparse MQA logits for the indexer and Mega-Gate, which combines the gate GEMM with expert selection.
Why it matters
Teams serving DeepSeek-family models on Blackwell (SM100) can expect performance gains just by upgrading. The actual size of the gains is not given in the sources.
Sources
- Official v0.31.0 vLLM