NaiveAI Open-Weights 309B Naive-N0.5-Flash With No Full Attention
TL;DR
- Naive-N0.5-Flash runs 39 sliding-window attention layers plus 9 DeepSeek Sparse Attention layers with zero full-attention layers across all 48 layers.
- The 309B parameter count masks 15.5B active parameters per token, keeping per-token compute closer to a mid-size dense model.
- NaiveRT peaks at 2,000 tokens per second in Ultrafast mode via mega-kernel fusion, Programmatic Dependent Launch, and FP8 speculative decoding — but RuntimeWire flags these as narrow decode-speed figures, not end-to-end benchmarks.
NaiveAI released Naive-N0.5-Flash, a 309B-parameter mixture-of-experts model with 15.5B active parameters, native 1M-token context, and no full-attention layers anywhere in the stack, published under an MIT license on Hugging Face.
The 48-layer transformer is arranged as eight six-layer modules, each pairing five sliding-window attention layers with one DeepSeek Sparse Attention layer, for 39 SWA and 9 DSA layers in total. The sliding window is 128 tokens; the sparse layer selects the top 2,048 tokens for backbone attention. Grouped-Query Attention uses 4 KV groups, and the DSA indexer runs 16 query heads. The model card describes the result as "Native 1M context without full attention through hybrid SWA–DSA architecture."
Training ran on 3.25T tokens at native 1M-token context, starting from Xiaomi's open-weight MiMo-V2.5 base: a 50B-token indexer warmup, 3T tokens of sparse-attention training, and a 200B-token learning-rate decay. The acknowledgments credit the Xiaomi MiMo team for the base model, the DeepSeek team for the sparse-attention design, and SGLang for inference infrastructure.
Serving runs on the team's own NaiveRT stack with mega-kernel fusion, Programmatic Dependent Launch, speculative decoding, and FP8 mixed precision. NaiveAI reports "Up to 2,000 tokens/s (Ultrafast mode)" and 50 tokens/s per user in a standard mode. The hosted API is priced at $0.10 per million input tokens and $0.40 per million output. The BibTeX title is "Naive-N0.5-Flash: Building Frontier AI with AI," and the card summarizes the target workload plainly: "Optimized for coding and AI R&D tasks."
It lands amid a busy week for open-weight releases in our open-source AI tracker, on the same day DeepSeek published its DSec elastic-compute paper.
What others are reporting
-
RuntimeWire Read →
Scrutinizes the AI-assisted development claim and contextualizes leadership (Jifeng Dai, ex-Microsoft Research Asia / SenseTime), flagging inference metrics as decode-speed-only rather than end-to-end figures.
AI systems can help build the next generation of AI
Originally reported by huggingface.co
Read the original article →Original headline: NaiveAI Open-Weights 309B MoE Naive-N0.5-Flash With No Full-Attention Layers, Trained by AI-Assisted Researchers