ByteDance pretrains 10-trillion parameter model, FT reports
TL;DR
- The FT sourced the 10T figure from three people familiar with the project; ByteDance has neither confirmed nor denied it.
- The model uses a Mixture-of-Experts architecture, meaning active parameters per token will be far below 10T, making raw comparisons to dense or competitor models structurally misleading.
- The training run requires roughly 30,000 GPUs over 3 to 6 months, putting a potential completion window in early-to-mid 2027.
A ten-trillion-parameter training run is a positioning move as much as an engineering one, and that is the frame worth putting on the Financial Times report, circulated this week, that ByteDance is pretraining an AI model with up to 10 trillion parameters. Three people familiar with the project told the FT the model is currently in pretraining, a phase that typically takes three to six months.
For scale, that would be roughly three times the size of Moonshot's Kimi K3, currently the largest Chinese model at around 2.8 trillion parameters, and it would put the TikTok parent company in the same ballpark as Anthropic's Mythos 5, which industry estimates place at around eight trillion parameters. Per The Decoder's writeup, founder Zhang Yiming told the 2,000-person Seed team internally to aim for world-leading model capabilities over the long term, and one source said ByteDance has avoided distillation, meaning training on outputs from other companies' models, for over a year. That matches Zhang's public no-distillation stance from last week.
The reason that matters is that Chinese frontier labs have spent the last year fighting off accusations that their gains came from copying Western model outputs rather than building from scratch. A ten-trillion-parameter run is expensive and slow, and it is also the cleanest possible answer to that accusation. It also plants ByteDance in the very small club of labs actively pretraining at frontier scale, alongside xAI, which is reportedly training Grok variants with six and ten trillion parameters on its Colossus 2 cluster, according to Elon Musk.
The FT's own reporting is thin on specifics: it does not name the model, does not disclose how many parameters would actually activate per query, does not specify the chips being used, and does not provide a release date. Total parameter counts are a ceiling on capability, not a description of it, and ByteDance has not shown its serving math. Any Chinese frontier training run also runs into the ongoing chip-export question the reporting sidesteps.
If the run lands, the interesting shift is on the distribution side. A frontier model paired with TikTok and the Doubao assistant, already among China's most-used AI products in the reporting, is a combination Western labs cannot easily match on reach, and it forces incumbents to price against a competitor whose consumer surface is already global.
What others are reporting
-
Reuters Read →
Reuters wire pick-up of the FT exclusive, adding the cross-model parameter comparison framework and a 3 to 6 month pretraining timeline as the primary development context.
With 10 trillion parameters, the model would be more than three times the size of Chinese startup Moonshot AI's Kimi K3.
-
XenoSpectrum Read →
Argues that MoE active-parameter counts make total-parameter comparisons structurally invalid, and that Anthropic's opacity makes competitive assessment impossible without matched third-party benchmarks.
10 trillion is not a finished performance benchmark—it is simply the upper bound of model scale currently under consideration.
-
Crypto Briefing Read →
Provides the most specific infrastructure detail: 30,000 GPUs, MoE architecture, 2,000-person Seed AI team, MegaScale systems designed for 10,000-plus-GPU runs, and US export-control context.
The effort would require approximately 30,000 GPUs and an estimated 3 to 6 months of continuous pre-training.
-
Slashdot Read →
Community discussion surfaces training-data scarcity as the likely binding constraint, arguing current frontier models are already undertrained relative to parameter count regardless of scale.
Current top models are already undertrained by multiple orders of magnitude vs their number of parameters...for lack of data.
Originally reported by thenextweb.com
Read the original article →Original headline: FT: ByteDance Pretraining a 10-Trillion-Parameter Model, More Than 3x the Size of Moonshot's Kimi K3