Chinese AI labs still train on Nvidia as CUDA lock-in bites
TL;DR
- An AI researcher at a Shanghai-affiliated institute estimates that switching training from Nvidia CUDA to Huawei Ascend would add at least 50% in time and costs.
- Training LLMs on Nvidia chips remains the norm among Chinese AI developers, an unnamed industry source told the South China Morning Post.
- The bottleneck is software, not silicon: existing CUDA pipelines cannot run directly on Ascend and require extensive rewriting on Huawei's CANN stack.
Chinese AI labs keep training their frontier models on Nvidia silicon, and the reason is not raw compute. It is code. The South China Morning Post reports that porting existing CUDA training pipelines to Huawei's Ascend chips and its CANN software stack would, on one researcher's estimate, add at least 50 per cent in time and cost.
The researcher, James Wang at a Shanghai-affiliated institute, put it plainly to the paper: "Our existing training pipelines are reliant on CUDA. CUDA code cannot run directly on Ascend and requires extensive rewriting." Another industry source told the SCMP that "training LLMs on Nvidia chips for now remains the norm among Chinese AI developers." The 50 per cent figure is one team's estimate, not an industry-wide benchmark, and it is worth reading that way.
That framing flips the louder narrative around China's chip push. Most of the year's coverage has been about hardware supply, export controls, and the capital being marshalled to fund domestic fabs. Our own tracker has logged 358 China AI stories in the last 90 days, and the recent ones sit mostly on that hardware-and-capital side, from Beijing's $28T capital-markets push to the BIS probe of offshore Nvidia access. But if the binding constraint at the top of the stack is software rather than silicon, more Ascend wafers do not close the gap on their own. CUDA's ecosystem lead is the moat, and it is a moat measured in developer-hours.
The SCMP does not name which specific Chinese labs are still on Nvidia versus which have migrated, and it does not clarify whether Wang's 50 per cent figure is a one-time port cost or a recurring training overhead. It also does not quantify how large Huawei's installed Ascend base at the top labs is today. Those blanks are worth holding in mind before treating one researcher's number as a national trend.
The path down for Beijing runs through tooling. If Huawei or Chinese cloud vendors can absorb enough of the CANN rewrite work through PyTorch-compatible backends, managed training environments, or reusable recipes, that premium can shrink over time. Whether the next crop of Chinese frontier models is trained on Ascend or Nvidia will say more about how far that tooling work has come than about any new fab announcement.
Originally reported by scmp.com
Read the original article →Original headline: SCMP: China's Top AI Labs Still Train on Nvidia Chips as Huawei CANN Switch Adds 50%+ in Cost