vLLM v0.31.0 (717 commits, 307 contributors, 96 new) adds the vllm preload weight-cache daemon for fast engine restarts and experimental CRIU engine snapshots, makes FlashMLA mega attention with NVFP4 compressed KV the SM100 default for DeepSeek-V4.1-Flash, brings draft-model speculative decoding to Model Runner V2, gates per-request multimodal kwargs behind a flag, and fixes prefix-cache key collisions between LoRA names and cache_salt. Several breaking changes ship too.
Key Takeaways
- ✓Scale: 717 commits from 307 contributors (96 new).
- ✓Fast restart:
vllm preloadkeeps post-quantized weights resident in GPU memory across restarts (DP, MTP drafts, /health); experimental CRIU snapshots restore an initialized TP1 engine. - ✓Security: per-request multimodal kwargs rejected unless
--trust-request-mm-kwargs; LoRA path now part of the block hash to stop prefix-cache collisions. - ✓Breaking:
tokenizer_mode="slow"removed; onlinequantization="fp8"replaced byfp8_per_tensor;--enforce-eageralso disables JIT warmup. - ✓Model paths: GLM-5.3-Flash metadata ops 1.6–4.8x faster and 3 GiB indexer workspace saved; fused KimiViT QK RoPE up to 29x; ~25 s faster Qwen3.8-Flash-Next weight loading on DGX Spark.
Developer 3s Key Decision Metrics
Not enough VRAM? Compare cloud API and self-hosting costs
Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
vLLM v0.31.0 (717 commits, 307 contributors) targets restart cost and multi-tenant safety. The new vllm preload daemon keeps post-quantized weights resident in GPU memory across engine restarts (with DP, MTP drafts, a /health endpoint and readiness wait), and experimental vllm snapshot create/restore uses CRIU to restore an initialized TP1 engine. DeepSeek-V4.1-Flash now defaults to FlashMLA mega attention with NVFP4 compressed KV on SM100; Model Runner V2 gains draft-model speculative decoding and custom logits processors; --max-num-active-seqs caps running admission. Security: per-request multimodal kwargs are rejected unless --trust-request-mm-kwargs is set, and LoRA names/paths can no longer collide with cache_salt in prefix-cache keys. No end-to-end throughput headline was published; component numbers include 1.6–4.8x faster GLM-5.3-Flash metadata ops, up to 29x fused KimiViT RoPE, and ~25 s faster Qwen3.8-Flash-Next loading on DGX Spark. Breaking changes: tokenizer_mode="slow" removed, online FP8 via fp8_per_tensor, renamed Mamba prefix-cache flag, --enforce-eager also disables JIT warmup. Install with pip install vllm or vllm/vllm-openai:v0.31.0.
Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.