← October 7, 2026 briefing

P2 Open source Generally available vLLM

vLLM v0.31.0 released, focused on DeepSeek-V4.1-Flash inference performance

Announced October 6, 2026

What happened

vLLM v0.31.0 includes 717 commits from 307 contributors (96 of them new). The main change is better DeepSeek-V4.1-Flash performance: FlashMLA mega attention with V4.1 NVFP4 compressed KV cache is now the default on SM100. It also adds kernel fusions such as DeepGEMM sparse MQA logits for the indexer and Mega-Gate, which combines the gate GEMM with expert selection.

Why it matters

Teams serving DeepSeek-family models on Blackwell (SM100) can expect performance gains just by upgrading. The actual size of the gains is not given in the sources.

Sources