theregister.com web signal

AMD debuts Threadripper Halo workstation with 576GB HBM3e

AMD Chips AI Infrastructure Inference ai-infrastructure

TL;DR

  • The Threadripper Halo pairs a 96-core Threadripper PRO 9995WX with up to four MI350P Instinct accelerators for 576 GB of HBM3e and 16 TB/s of bandwidth.
  • AMD is pitching it as capable of running trillion-parameter models at four-bit precision entirely in GPU memory, with system RAM offload for larger models.
  • The workstation is set to launch in 2027 at an expected $100,000 to $150,000, aimed against Nvidia's DGX Station.

AMD used its IFA 2026 slot to introduce the Threadripper Halo, a liquid-cooled workstation topping out at 576 GB of HBM3e and 16 TB/s of memory bandwidth. It is set to launch in 2027.

The system pairs a 96-core Threadripper PRO 9995WX with up to four PCIe MI350P Instinct accelerators, each carrying 144 GB of HBM3e at 4 TB/s, backed by up to 2 TB of DDR5 for 2.6 TB of combined system memory. It is designed to run "models exceeding a trillion parameters in size (at four-bit precision)" entirely in GPU memory, with system RAM offload for larger models like Moonshot.AI's 2.8 trillion-parameter Kimi K3, The Register reports.

AMD is aiming the box straight at Nvidia's DGX Station, claiming "3.4x the total system memory and more than twice the memory bandwidth" of the rival. The Register's Tobias Mann is unflattering about the assembly, writing that AMD "just raided its spare parts bin and cobbled together a DGX Station rival"; the Threadripper PRO 9995WX debuted in 2025 and the MI350P is an existing accelerator, so nothing here is a new silicon reveal.

A practical wrinkle: the IFA demo units held only two GPUs. A fully populated quad-MI350P configuration would exceed a standard North American outlet without an electrical upgrade, or a step down from the 600W GPU profile to the more conservative 450W. AMD has not set a price; the article cites an expected range of $100,000 to $150,000. It joins a busy run of AI infrastructure coverage on the tracker as spending reshapes around ever-larger local inference, alongside work on faster serving like random attention matching top KV scorers in vLLM.