Open-Source Strata Engine Runs 125B Qwen3.8-Flash-Next on a Single 12GB Consumer GPU at 60-95 Tokens/Sec
Summary
Strata is an open-source inference engine letting Qwen3.8-Flash-Next, a 125B-parameter MoE model, run on a single NVIDIA or AMD gaming GPU with 12GB+ VRAM and 64GB RAM, writing at 60-95 tokens per second. It uses a specialist-routing scheme where only 10 of 24,576 'experts' fire per token, plus a built-in speculative drafter claimed to deliver 1.6-1.8x speedup over conventional inference. The project reached the Hacker News front page on October 4 with 72 points.
Originally reported by github.com
Read the original article →Original headline: Open-Source Strata Engine Runs 125B Qwen3.8-Flash-Next on a Single 12GB Consumer GPU at 60-95 Tokens/Sec