huggingface.co web signal

Qwen sets Aug 26 drop for Qwen3.8-Flash-Next, a Qwen4 preview

TL;DR

  • Alibaba's Qwen team put a countdown on Hugging Face for Qwen3.8-Flash-Next, billed as a preview of the Qwen4 architecture, with weights due August 26, 2026.
  • Community screenshots of a ModelScope listing describe the model as a 125B-parameter MoE with 51B additional N-gram embeddings and 6B parameters active per token; Qwen has not officially confirmed the numbers.
  • A preview writeup reads the Qwen4 stack as a Gated Delta Network linear-attention layer paired with a still-undocumented mechanism called Qwen Sparse Attention.

Alibaba's Qwen team posted a countdown page on Hugging Face for Qwen3.8-Flash-Next, a model the listing describes only as "A Preview of the Qwen4 Architecture," with weights due August 26, 2026.

The specs circulating come from community screenshots of a ModelScope listing rather than an official Qwen document. "Qwen3.8-Flash-Next just appeared on ModelScope. 125B MoE + an additional 51B of N-gram embeddings, with only 6B parameters activated," Neo wrote on X. Daniel Han of Unsloth AI posted that it is "a new open-weight multimodal MoE model" and said his team is "working on @UnslothAI day zero support."

The landing page itself publishes no benchmarks, no license, no parameter count and no context window; it lists 595 users signed up for release notifications. A preview writeup by Orca Router reads the Qwen4 architecture as pairing a linear-attention layer called Gated Delta Network with a still-undocumented mechanism named Qwen Sparse Attention.

Shared on Bluesky by 1 AI expert