github.com via Hacker News

Strata Runs 125B Qwen3.8-Flash-Next on a 12GB Gaming GPU

TL;DR

  • Open-source Strata runs Qwen3.8-Flash-Next's 125 billion parameters on a single consumer GPU with 12 GB of VRAM and 32 GB of system RAM.
  • The model activates only 10 of 24,576 experts per token, and a built-in speculative drafter is claimed to add a 1.6-1.8x throughput gain.
  • Self-reported figures include 94 tokens per second on an NVIDIA RTX 5070 and 60 on an AMD RX 9070 XT at Q2_0 quantization.

Strata, an MIT-licensed open-source inference engine, claims to run Qwen3.8-Flash-Next, a 125-billion-parameter model, on a single consumer GPU with 12 GB of VRAM and 32 GB of system memory.

The README's architectural explanation is unusually plain: "The model is a team of 24,576 small specialists ('experts'). Each word needs only 10 of them." That sparse routing is what lets a nominally 125B model fit a gaming rig at all. Strata keeps the hot experts on the GPU and offloads the rest to system RAM, which the project describes in kitchen terms: "the things you use all the time stay on the counter, and the rest waits in the pantry."

On an NVIDIA RTX 5070, Strata's own figures report 94 tokens per second on output and 2,650 tokens per second on input at Q2_0 quantization. On an AMD RX 9070 XT, the same quantization yields 60 output and 1,160 input tokens per second. A speculative drafter is claimed to add another 1.6-1.8x on top: "A small helper guesses the next few words. The big model checks them all at once." The IQ3_S variant on the same NVIDIA card drops to 53 tokens per second output, trading throughput for quality.

The weights were compressed by ISTA-DASLab, UkisAI, and Unsloth; the engine credits llama.cpp/ggml. No third-party benchmarks are posted yet, so every figure here is the project's own claim until someone reproduces it. Strata lands in a visible run of local-inference releases on our tracker, after Cloudflare's Clef earlier this week and AI2's Olmo-core 3 MoE throughput work yesterday.