Shanghai AI Lab’s LMDeploy shipped v0.18.0 on 2026-09-28: input logprobs, expanded SM90 quantized GEMMs with unified TurboMind linear paths, Ascend GLM-5.2, Qwen3.5 dflash, and MoE shared-expert/FFN sharding. PyTorch path prototypes TurboMind W4A16 (AWQ), piecewise CUDA Graph prefill, XTuner TileLang sparse MLA, checkpoint-engine weight updates, and request-only KV cache metrics.

Key Takeaways

  • ✓Shipped v0.18.0 — Qwen3.5 dflash, Ascend GLM-5.2, SM90 quantized GEMMs
  • ✓TurboMind W4A16 AWQ prototype + piecewise CUDA Graph prefill + TileLang sparse MLA
  • ✓checkpoint-engine weight updates + request-only KV cache metrics + input logprobs
  • ✓Upgrade: pip install -U "lmdeploy>=0.18.0"
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

Core Background & Industry Pain Points

Private LLM serving must squeeze Hopper (SM90), Ascend, and MoE hybrids. Pain points: lagging paths for Qwen3.5/GLM-5.2; fragmented quantized GEMM/linear kernels; hard to combine AWQ W4A16 with CUDA Graph prefill; weak weight hot-update and per-request KV observability.

Architecture Highlights & Internals

v0.18.0 adds input logprobs, expands SM90 quantized GEMMs, Ascend GLM-5.2, Qwen3.5 dflash, and MoE shared-expert/FFN sharding. PyTorch prototypes TurboMind W4A16 AWQ, piecewise CUDA Graph prefill, TileLang sparse MLA, checkpoint-engine weight updates, and request-only cache metrics. TurboMind moves to C++20; tool params can stream incrementally.

Authoritative Benchmarks & Measured Scores

No unified tokens/s table in the notes (docs only fix A100 FP16 labels). Benchmark Qwen3.5 dflash, W4A16, and piecewise CUDA Graph on your GPUs; do not generalize PR wording into cross-model speedups.

Developer Hands-on Guide

pip install -U "lmdeploy>=0.18.0". Follow docs; validate dflash/W4A16/CUDA Graph and tool_choice before production. See the v0.18.0 release.