llama.cpp v0.6.0 Adds GLM-5.3-Flash, 3x Apple Metal Matmul
TL;DR
- llama.cpp v0.6.0 ships day-one support for GLM-5.3-Flash, which the project describes as a 320B text+vision hybrid model.
- MTP speculative decoding for Qwen4Exp delivers a ~1.5x decode speedup on DGX Spark, according to the release notes.
- New Metal MMA kernels claim up to ~3x faster mat-mul on Apple GPUs for speculative and batched decoding.
The llama.cpp v0.6.0 release notes advertise day-one support for GLM-5.3-Flash, which the project calls "a 320B text+vision hybrid model," and MTP speculative decoding for Qwen4Exp credited with a "~1.5x decode speedup on DGX Spark."
Apple Silicon gets the hardware headline. The notes describe "new few-row MMA mat-mul kernels for speculative and batched decoding, up to ~3x faster mat-mul on Apple GPUs," alongside a Metal tensor-API flash attention kernel for F16 KV. Sparse flash attention kernels also land on SYCL, Vulkan and Metal in the underlying ggml bump to v0.26.0, which adds a "model-driven W4A4 (NVFP4/MXFP4) mul_mat path."
On the server side, a new `/v1/systemone` API arrives "supporting five decision models - laya, julia-1, lev, openjev (+vision), kev," alongside Clef, Nimble and Ling 3.0 VL support. Session formats bump to LLAMA_SESSION_VERSION 11 and LLAMA_STATE_SEQ_VERSION 4. It joins a dense run of local-inference releases in our open-source feed, which has logged 312 stories in the past 90 days.
Originally reported by github.com
Read the original article →Original headline: llama.cpp v0.6.0 Lands MTP Speculative Decoding for Qwen4Exp and First-Party GLM-5.3-Flash, Clef and Ling 3.0 VL Support