interconnects.ai web signal

Kimi K3, Qwen 3.8 push China's open-weight AI escalation

TL;DR

  • Moonshot's Kimi K3 and Alibaba's 2.4 trillion-parameter Qwen 3.8 preview arrived within days, both aimed at frontier-tier open weights.
  • Nathan Lambert pushed back on Ben Thompson's Stratechery claim that distillation drives Chinese open models, arguing it matters less as training shifts to RL.
  • Xi Jinping used his WAIC keynote to commit China's AI ecosystem to open source, framing open weights as a global-diffusion strategy.

Nathan Lambert's latest Interconnects recap with Florian Brand picks up where the flagship benchmark drama usually ends and asks a different question: what happens now that a handful of Chinese labs are shipping open-weight models at frontier scale on a one-to-two month cadence? The immediate examples are Moonshot's Kimi K3, described as a 2.8 trillion-parameter open-weight release, and Alibaba's Qwen 3.8, previewed at 2.4 trillion parameters and, unusually for the largest Qwen tier, promised as an eventual open-weight drop.

Most of the conversation is about what these models feel like to use rather than leaderboard placement. Brand puts Kimi K3's coding at roughly a "54-55 level" versus tools like Codex, good enough for real workflow tasks, still a step behind the closed frontier on the harder edges. Lambert says Kimi surprised him on research-style queries, and both flag that the more interesting story is release cadence: DeepSeek, Zhipu's GLM, Qwen, MiniMax and Moonshot are now iterating fast enough that the frontier tier and the near-frontier tier keep swapping places on specific tasks.

The recap also pushes back on Ben Thompson's recent Stratechery argument that distillation from closed frontier models does most of the heavy lifting behind Chinese open-weight releases. Lambert's counter is that as training shifts from SFT toward reinforcement learning, using GPT or Claude as graders across millions of RL rollouts would be prohibitively expensive; distillation, he says, "is getting less and less impactful," and treating it as the main story risks pulling in regulation aimed at the wrong problem. Sitting behind all of this is Xi Jinping's WAIC keynote, which Lambert reads as a state-level commitment to open source as an export strategy, a very different backdrop from the US ecosystem where Thinking Machines, Nvidia, Poolside and Reflection are the labs being watched as counterweights.

The honest caveat is that the parameter counts here are total-model figures, not active-per-token compute. As MarkTechPost noted on the Qwen 3.8 preview, the active-parameter count is "the number nobody has," so real serving costs and deployment feasibility are still unknown, and the benchmarks published so far are self-reported. What the recap does not give you is a firm date or license for when Qwen 3.8's weights actually ship, or a clear read on whether any US challenger closes ground before year-end.

If Lambert is even directionally right, the shift worth planning around is not a single leaderboard flip but a bifurcation: the very hardest capabilities stay behind closed APIs, while near-frontier work, day-to-day coding, agentic tool use, research triage, becomes something anyone with a GPU can run.

Shared on Bluesky by 1 AI expert