llama.cpp Merges 320B GLM-5.3-Flash Hybrid Text-Vision Support
TL;DR
- GLM-5.3-Flash, a 320B hybrid text-and-vision model, was merged into llama.cpp on September 30, 2026 via PR #27773 by contributor timkhronos.
- The build mixes 34 KDA linear layers (reused from Kimi-K3) with 11 Dynamic Sparse Attention layers over an MLA-only attention path.
- Logits match the HuggingFace transformers reference on prefill and single-token decode; MTP support is deferred to follow-on PR #27917.
A 320-billion-parameter hybrid text-and-vision model, GLM-5.3-Flash, now runs natively in llama.cpp. The pull request from contributor timkhronos was merged on September 30, stitching together 34 KDA linear layers with 11 DSA (Dynamic Sparse Attention) layers on top of an MLA-only attention path.
The implementation leans on prior work. KDA layers reuse the Kimi-K3 code, the mHC controller is adapted from Deepseek V4, and the MoE with clamped SwiGLU follows the same DeepSeek lineage.
The new piece is an indexer that, in the author's words, 'scores pools of 4 consecutive token, and always keeps the incomplete tail'. A `llama_memory_hybrid_dsa` cache handles the hybrid recurrent-plus-sparse state, and the vision tower encoder ships with per-head QK-norm and optional image-token budgeting.
Correctness was checked against the Hugging Face reference. The author reports that 'logits match transformers on a small random model across full prefill, small ubatches and single token decode while sparse selection is active', with vision embeddings agreeing to roughly 1e-5. About 1GB of precision-sensitive tensors, including the indexer and KDA gates, are kept unquantized.
Multi-token prediction is deferred to a follow-on PR (#27917), and multi-sequence runs require the `--kv-unified` flag. Pre-converted GGUF quantizations are already up on Hugging Face in the avar6 repository. The merge lands within days of Anthropic's report that the GLM-5.3 family matches Claude on autonomous exploits, giving local runners a fast route to a frontier-class Chinese open-weight model.
Originally reported by github.com
Read the original article →Original headline: llama.cpp Merges Support for Zhipu's 320B Hybrid GLM-5.3-Flash With MoE, MLA, DSA and Vision