ByteDance pretrains 10-trillion parameter model, FT reports
TL;DR
- ByteDance is pretraining an AI model with up to 10 trillion parameters, per the Financial Times, citing three people familiar with the project.
- That would be roughly three times Moonshot's Kimi K3 at 2.8 trillion parameters and near Anthropic's Mythos 5, estimated around 8 trillion.
- Founder Zhang Yiming told the 2,000-person Seed team to aim for world-leading capabilities and avoid training on rival models' outputs.
A ten-trillion-parameter training run is a positioning move as much as an engineering one, and that is the frame worth putting on the Financial Times report, circulated this week, that ByteDance is pretraining an AI model with up to 10 trillion parameters. Three people familiar with the project told the FT the model is currently in pretraining, a phase that typically takes three to six months.
For scale, that would be roughly three times the size of Moonshot's Kimi K3, currently the largest Chinese model at around 2.8 trillion parameters, and it would put the TikTok parent company in the same ballpark as Anthropic's Mythos 5, which industry estimates place at around eight trillion parameters. Per The Decoder's writeup, founder Zhang Yiming told the 2,000-person Seed team internally to aim for world-leading model capabilities over the long term, and one source said ByteDance has avoided distillation, meaning training on outputs from other companies' models, for over a year.
The reason that matters is that Chinese frontier labs have spent the last year fighting off accusations that their gains came from copying Western model outputs rather than building from scratch. A ten-trillion-parameter run is expensive and slow, and it is also the cleanest possible answer to that accusation. It also plants ByteDance in the very small club of labs actively pretraining at frontier scale, alongside xAI, which is reportedly training Grok variants with six and ten trillion parameters on its Colossus 2 cluster, according to Elon Musk.
The honest caveat is that the FT report does not name the model, does not disclose how many parameters would actually activate per query, does not specify the chips being used, and does not provide a release date. Total parameter counts are a ceiling on capability, not a description of it, and ByteDance has not shown its serving math. Any Chinese frontier training run also runs into the ongoing chip-export question the reporting sidesteps.
If the run lands, the interesting shift is on the distribution side. A frontier model paired with TikTok and the Doubao assistant, already among China's most-used AI products in the reporting, is a combination Western labs cannot easily match on reach, and it forces incumbents to price against a competitor whose consumer surface is already global.
Originally reported by thenextweb.com
Read the original article →Original headline: FT: ByteDance Pretraining a 10-Trillion-Parameter Model, More Than 3x the Size of Moonshot's Kimi K3