theinformation.com web signal

Nvidia Tests Lower-Memory Rubin Ultra Variants Amid HBM Squeeze

TL;DR

  • Nvidia is evaluating three alternatives to Rubin Ultra's original 12-Hi HBM4e design: 8-Hi HBM4e, 12-Hi HBM4, and 8-Hi HBM4.
  • The mainline SKU previewed to customers reportedly drops to 8-Hi and 192GB, below Rubin's 12-Hi 288GB, while keeping peak FLOPS.
  • TrendForce attributes the retreat to tight DRAM supply through 2027, uncertain 12-Hi HBM4e validation, and data-center power constraints.

Nvidia has stopped treating the Rubin Ultra's memory spec as fixed. The Information reported that the company is weighing lower-memory versions of its next flagship training chip because it may not secure enough HBM. Corroborating analysis from TrendForce says Nvidia has begun evaluating three alternatives beyond the original 12-Hi HBM4e design: 8-Hi HBM4e, 12-Hi HBM4, and 8-Hi HBM4.

The direction of travel is a step down. The mainline Rubin Ultra SKU being previewed to customers reportedly keeps its peak theoretical FLOPS on HBM4, but memory drops to 8-Hi and 192GB, below Rubin's 12-Hi 288GB. That is an unusual retreat for a flagship AI accelerator, where more memory per package has been the default assumption of every roadmap slide for two years. It also fits a longer arc of Rubin Ultra spec cuts, from the 4-die 1TB HBM4E 16-Hi configuration previewed at GTC 2025, to HBM4E 12-Hi, to a canceled 4-die MCM, to the current 2-die package.

Why this matters for anyone building training clusters: memory per GPU caps how much of a model each node can hold, and buyers who planned around 288GB per Rubin Ultra will now need to re-plan around roughly two-thirds of that. TrendForce attributes the retreat to tight DRAM supply through 2027, uncertain validation timelines for 12-Hi HBM4e, and data-center power shortages, not to any problem with the silicon itself. On the power side, an 1800W Max-Q is expected to be the mainstream variant, with a 1200W SKU aimed at lower-compute workloads such as token decoding and a 2600 to 2800W Max-P option at the top.

The honest caveat is that these are described as options Nvidia is evaluating with no final decision yet reached, and the final SKU mix could still change. What the reporting does not give you is the breakdown between how much of the cut is HBM supply, how much is power, and how much is cost, or which hyperscaler customers have signed off on the lower spec.

If the mainline product does land at 192GB, the interesting knock-on is who wins on the margins. Memory suppliers that clear 12-Hi HBM4e validation first still have a shot at the higher-tier SKU, and buyers whose real workload is token decoding rather than largest-possible-model training may find the 1200W part fits them better than the flagship anyway.