huggingface.co web signal

NaiveAI Open-Weights 309B Naive-N0.5-Flash With No Full Attention

TL;DR

  • NaiveAI shipped Naive-N0.5-Flash under MIT: 309B total parameters, 15.5B active, native 1M-token context, no full-attention layers.
  • The 48-layer stack pairs 39 sliding-window attention layers with 9 DeepSeek Sparse Attention layers, built on Xiaomi's open MiMo-V2.5 base.
  • NaiveRT inference claims up to 2,000 tokens/s in Ultrafast mode; hosted API lists $0.10 input and $0.40 output per million tokens.

NaiveAI released Naive-N0.5-Flash, a 309B-parameter mixture-of-experts model with 15.5B active parameters, native 1M-token context, and no full-attention layers anywhere in the stack, published under an MIT license on Hugging Face.

The 48-layer transformer is arranged as eight six-layer modules, each pairing five sliding-window attention layers with one DeepSeek Sparse Attention layer, for 39 SWA and 9 DSA layers in total. The sliding window is 128 tokens; the sparse layer selects the top 2,048 tokens for backbone attention. Grouped-Query Attention uses 4 KV groups, and the DSA indexer runs 16 query heads. The model card describes the result as "Native 1M context without full attention through hybrid SWA–DSA architecture."

Training ran on 3.25T tokens at native 1M-token context, starting from Xiaomi's open-weight MiMo-V2.5 base: a 50B-token indexer warmup, 3T tokens of sparse-attention training, and a 200B-token learning-rate decay. The acknowledgments credit the Xiaomi MiMo team for the base model, the DeepSeek team for the sparse-attention design, and SGLang for inference infrastructure.

Serving runs on the team's own NaiveRT stack with mega-kernel fusion, Programmatic Dependent Launch, speculative decoding, and FP8 mixed precision. NaiveAI reports "Up to 2,000 tokens/s (Ultrafast mode)" and 50 tokens/s per user in a standard mode. The hosted API is priced at $0.10 per million input tokens and $0.40 per million output. The BibTeX title is "Naive-N0.5-Flash: Building Frontier AI with AI," and the card summarizes the target workload plainly: "Optimized for coding and AI R&D tasks."

It lands amid a busy week for open-weight releases in our open-source AI tracker, on the same day DeepSeek published its DSec elastic-compute paper.