huggingface.co web signal

XPeng AI's X-AuT Trims Speech-LLM Audio Layers Without Loss

TL;DR

  • Pruning Qwen3-ASR-0.6B from 18 to 16 layers cut mean error from 5.61% to 5.27% across ten Chinese-English benchmarks.
  • A 14-layer variant holds error at 5.75% while removing 20.7% of audio-tower parameters.
  • Progressive 18→14 pruning beats a direct 18→14 cut, 5.75% versus 6.73% mean error.

XPeng AI's X-AuT framework prunes audio-encoder layers from the Qwen3-ASR-0.6B speech model and, going from 18 to 16 layers, cuts mean error across ten Chinese-English benchmarks from 5.61% to 5.27%, according to a paper posted to Hugging Face by Shiyu Huang and colleagues at XPENG AI.

The more aggressive 14-layer variant lands at "5.75% error with 20.7% fewer audio-tower parameters", the authors report, and a progressive 18→14 schedule beats a direct 18→14 prune at 6.73%. A 1.7B teacher model registers 5.55% mean error against 8.45% for self-distillation, marking the gap the cross-scale supervision is bridging.

The recipe combines behavioral probes for layer selection, representation alignment, cross-scale distillation and LoRA adaptation with a frozen language-model backbone. Weights and code are posted at the XPENG-AI/X-AuT repository, joining a busy quarter of voice-AI research our tracker has logged. Recent open speech work in that log includes AuK's 1.5B model unifying generation and editing, one of 42 voice-AI stories over the last 90 days.