artificialanalysis.ai via Reddit

Artificial Analysis benchmark: MoE active-param count beats GPU on local agents

TL;DR

  • Qwen3.6-35B-A3B, a 3B-active MoE, ran 2.5-3.3x faster than dense Qwen3.8-27B on every laptop and workstation tested.
  • NVIDIA's GeForce RTX 5090 finished 3.5x faster overall and over 5x faster on decode than 128 GB unified-memory rigs, tracing to 1,792 GB/s bandwidth.
  • Even with 73-93% of prompt tokens served from KV cache, prefill still consumed 22-41% of end-to-end completion time on Qwen3.8-27B.

Artificial Analysis has published results from AA-AgentPerf-Local, an open-source tool that "replays real agent trajectories on laptop & workstation hardware to test inference performance," and the finding across four rigs is that active parameter count dominates local-agent latency. Qwen3.6-35B-A3B, a mixture-of-experts model with 3B active parameters, ran fastest on every system tested and beat the dense Qwen3.8-27B by 2.5 to 3.3x.

The four systems measured were the NVIDIA GeForce RTX 5090 (32 GB), NVIDIA DGX Spark and AMD Ryzen AI Halo (both 128 GB), and the 64 GB MacBook Pro M5 Pro, all running 4-bit quantized models on a workload of 8 recorded agent tasks spanning 168 model turns. The 5090 finished 3.5x faster than the unified-memory systems on compatible models, and its decode ran more than 5x faster; Artificial Analysis attributes the gap to memory bandwidth, 1,792 GB/s on the 5090 versus 256-307 GB/s on the unified-memory devices. The three unified-memory devices produced similar decode speeds to each other, and the MacBook Pro M5 Pro landed within 2-9% of Ryzen AI Halo on two models while trailing by 21% on Qwen3.5-9B.

Even in a workload that served 73-93% of prompt tokens from key-value cache, prefill still accounted for 22-41% of end-to-end completion time on Qwen3.8-27B.

The user-facing consequence shows up as tool-return waits: the Ryzen AI Halo paused up to 24 seconds waiting on a tool return, the MacBook Pro up to 39 seconds, and the 5090 around 2 seconds. At launch MSRP the MacBook Pro M5 Pro (64 GB) was $3,700, and the DGX Spark and Ryzen AI Halo shipped at $4,000 each. Artificial Analysis says the next round will add x86 laptops, RTX PRO 6000 Blackwell cards and user-submitted leaderboards, and the tool is part of a busy week for local-inference coverage on our inference tracker.