Shanghai AI Lab’s LMDeploy shipped v0.18.0 on 2026-09-28: input logprobs, expanded SM90 quantized GEMMs with unified TurboMind linear paths, Ascend GLM-5.2, Qwen3.5 dflash, and MoE shared-expert/FFN sharding. PyTorch path prototypes TurboMind W4A16 (AWQ), piecewise CUDA Graph prefill, XTuner TileLang sparse MLA, checkpoint-engine weight updates, and request-only KV cache metrics.
Key Takeaways
- ✓Shipped v0.18.0 — Qwen3.5 dflash, Ascend GLM-5.2, SM90 quantized GEMMs
- ✓TurboMind W4A16 AWQ prototype + piecewise CUDA Graph prefill + TileLang sparse MLA
- ✓checkpoint-engine weight updates + request-only KV cache metrics + input logprobs
- ✓Upgrade: pip install -U "lmdeploy>=0.18.0"
Finished reading? Explore benchmark rankings & pricing
Real-world SWE-bench scores & $20/mo vs API cost break-even calculator
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
Core Background & Industry Pain Points
Private LLM serving must squeeze Hopper (SM90), Ascend, and MoE hybrids. Pain points: lagging paths for Qwen3.5/GLM-5.2; fragmented quantized GEMM/linear kernels; hard to combine AWQ W4A16 with CUDA Graph prefill; weak weight hot-update and per-request KV observability.
Architecture Highlights & Internals
v0.18.0 adds input logprobs, expands SM90 quantized GEMMs, Ascend GLM-5.2, Qwen3.5 dflash, and MoE shared-expert/FFN sharding. PyTorch prototypes TurboMind W4A16 AWQ, piecewise CUDA Graph prefill, TileLang sparse MLA, checkpoint-engine weight updates, and request-only cache metrics. TurboMind moves to C++20; tool params can stream incrementally.
Authoritative Benchmarks & Measured Scores
No unified tokens/s table in the notes (docs only fix A100 FP16 labels). Benchmark Qwen3.5 dflash, W4A16, and piecewise CUDA Graph on your GPUs; do not generalize PR wording into cross-model speedups.
Developer Hands-on Guide
pip install -U "lmdeploy>=0.18.0". Follow docs; validate dflash/W4A16/CUDA Graph and tool_choice before production. See the v0.18.0 release.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.