LMDeploy 0.18.0: Qwen3.5 dflash, SM90 quantized GEMMs, Ascend GLM-5.2; TurboMind W4A16 + piecewise CUDA Graph
Shanghai AI Lab’s LMDeploy shipped v0.18.0 on 2026-09-28: input logprobs, expanded SM90 quantized GEMMs with unified TurboMind linear paths, Ascend GLM-5.2, Qwen3.5 dflash, and MoE shared-expert/FFN sharding. PyTorch path prototypes TurboMind W4A16 (AWQ), piecewise CUDA Graph prefill, XTuner TileLang sparse MLA, checkpoint-engine weight updates, and request-only KV cache metrics.
- •Shipped v0.18.0 — Qwen3.5 dflash, Ascend GLM-5.2, SM90 quantized GEMMs
- •TurboMind W4A16 AWQ prototype + piecewise CUDA Graph prefill + TileLang sparse MLA
- •checkpoint-engine weight updates + request-only KV cache metrics + input logprobs
- •Upgrade: pip install -U "lmdeploy>=0.18.0"
