ggml-org released llama.cpp v0.6.0 on Oct 5 (v0.5.0 was Sep 23). It brings GLM-5.3-Flash (GLM5-Next, a 320B text+vision hybrid MoE) into a stable tag, adds MTP speculative decoding for Qwen4Exp (~1.5x decode on DGX Spark), introduces the llama_batch_ext / llama_process() API for mixed token and embedding batches, adds few-row MMA mat-mul kernels on Metal (up to ~3x for speculative and batched decoding), and bumps ggml to v0.26.0. Session file formats are bumped too.
Key Takeaways
- ✓New models: GLM-5.3-Flash (GLM5-Next, 320B KDA/DSA hybrid text+vision MoE), Clef decision model, Ling 3.0 VL, LFM2.5-Encoder
- ✓Speed: Qwen4Exp MTP speculative decoding ~1.5x decode on DGX Spark; Metal few-row MMA mat-mul up to ~3x
- ✓API: llama_batch_ext + llama_process() for mixed token/embedding batches; server and examples migrated
- ✓Compat: LLAMA_SESSION_VERSION 11 and LLAMA_STATE_SEQ_VERSION 4, so old session/state files must be regenerated
- ✓Server: /v1/models reports modalities, /v1/embeddings takes typed vision/audio/video input, new /v1/systemone endpoint
Developer 3s Key Decision Metrics
Turn your technical choice into a development budget
Compare 29+ dev plans & simulate token costs vs $20/mo subscriptions
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
llama.cpp v0.6.0 is the first stable tag since v0.5.0 (Sep 23) and brings GLM-5.3-Flash (GLM5-Next, a 320B KDA/DSA hybrid text+vision MoE, #27773) into a release, along with the Clef decision model, Ling 3.0 VL and LFM2.5 encoders. The main core change is the new llama_batch_ext API with llama_process() (#24669), which lets one batch mix raw tokens and embeddings and carry per-token state embeddings for MTP and deepstack models; server, mtmd and speculative decoding are migrated to it. There is also a model-driven W4A4 (NVFP4/MXFP4) mat-mul path, and ggml moves to v0.26.0.
The release notes report two speedups: Qwen4Exp MTP speculative decoding at about 1.5x decode on DGX Spark, and new few-row MMA mat-mul kernels on Metal at up to about 3x for speculative and batched decoding (#29869). No GLM-5.3-Flash throughput numbers are given. Before upgrading, note the session format bump (LLAMA_SESSION_VERSION 11, LLAMA_STATE_SEQ_VERSION 4), so saved sessions must be regenerated, and C API users should move to llama_batch_ext. Server's /v1/models now reports modalities and /v1/embeddings accepts typed vision/audio/video input.
Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.