Ten trillion parameters. That’s the number the Financial Times attached to ByteDance’s next AI model this week — and if it holds up, it would make the system more than three times larger than Kimi K3, currently one of the biggest open-weight models to come out of China. The headline number is eye-catching. Whether it means anything is a different question entirely.
According to the FT, which cited three people familiar with the project, ByteDance is training a model with as many as 10 trillion parameters, explicitly positioned to close the gap with Anthropic’s most capable frontier system — reportedly referred to internally as Mythos. That’s a specific, ambitious target, not a vague “we’re working on something big.” And it’s worth unpacking, because the number itself tells you almost nothing about whether ByteDance is actually about to leapfrog the field.
What ByteDance Is Actually Building
Strip away the headline figure and here’s what’s confirmed: the model is in early pre-training, and pre-training alone typically takes three to six months for something at this scale. Fine-tuning and safety work come after that. In plain terms — nothing is shipping soon. If the timeline holds, results (or complications) probably won’t surface publicly before the tail end of 2026, and a usable product is further out still.
For context, Moonshot AI’s Kimi K3 — the Chinese model most people currently point to as the country’s strongest open-weight release — sits at roughly 2.8 trillion parameters and was trained on around 20,000 Nvidia chips. Separate reporting on ByteDance’s project puts its compute footprint at closer to 30,000 GPUs, which tracks: a model more than three times the size needs meaningfully more silicon, not just more time. If ByteDance is genuinely targeting 10 trillion parameters, the compute bill alone puts this in a different league than anything else coming out of a Chinese lab so far. We covered Kimi K3’s release in detail when the weights dropped, and it’s a useful baseline for just how fast the scale is moving — Kimi K3 was the biggest open-weight Chinese model on the market a matter of weeks ago, and it’s already being framed as the small comparison point.

The Model Everyone’s Chasing
The “Mythos” framing is doing a lot of work here, and it’s worth being precise about what we actually know. The Financial Times describes it as one of the frontier systems Chinese AI labs have so far struggled to match — that’s the FT’s characterization of the competitive landscape, not a confirmed product name from Anthropic itself. Anthropic hasn’t put out a press release about a system called Mythos. What’s real is the underlying dynamic: every major Chinese AI lab is currently measuring itself against a small handful of American frontier models, and closing that gap has become the industry’s central obsession.
That’s the same dynamic behind the pricing war we wrote about when Chinese labs started undercutting OpenAI and Anthropic on cost — DeepSeek, Kimi, and Qwen aren’t just competing on capability anymore, they’re competing on capability per dollar. A model this size is a different kind of bet: it says ByteDance thinks the fastest way to close the gap isn’t a cheaper model, it’s a bigger one.
Here’s the Part Nobody’s Advertising: Size Isn’t the Whole Story
This is the detail that gets buried under the “10 trillion” headline, and it’s the one that actually matters if you want to know whether this model will be any good: parameter count alone is a weak predictor of real-world performance. None of the reporting on ByteDance’s project discloses whether this is a dense model or a Mixture-of-Experts (MoE) architecture — and that distinction is enormous. A 10-trillion-parameter MoE model might only activate a fraction of those parameters per query, making it far cheaper to run (and closer in practical terms to something like Kimi K3) than the raw number suggests. Without that split disclosed, “10 trillion” is closer to a marketing ceiling than a performance guarantee.
The industry has learned this lesson before, publicly and painfully. Every outlet covering this story landed on the same caveat: bigger models are not automatically better, and data quality, training technique, and post-training refinement often matter as much as raw scale. GPT-4.5 is the textbook example — a genuinely massive model that underwhelmed against smaller, more efficiently trained systems. Scale buys you a ceiling. It doesn’t buy you the floor.
Interesting Tangent: The No-Copying Order
One detail in the reporting is worth sitting with on its own. ByteDance founder Zhang Yiming reportedly told the team building this model to avoid AI distillation shortcuts — the practice of training a new model by learning from an existing frontier model’s outputs, which is faster and cheaper but caps how far ahead of the source model you can ever get. He wants genuine capability, not a copy. That’s a notable instruction, because distillation has quietly become one of the fastest ways for a lab with less raw compute to catch up to the frontier — and ByteDance is explicitly choosing the slower, more expensive, more original path instead. It’s a bet that paying full price for scale beats paying a discount for a ceiling.
Why This Matters Beyond One Company
Zoom out and this story isn’t really about ByteDance. It’s the latest data point in a race where “who can hire the people who know how to train these things” matters just as much as who has the most GPUs — something we dug into when we looked at why top AI researchers keep jumping between labs. Training a coherent, well-optimized 10-trillion-parameter model isn’t just a hardware problem; it’s an execution problem, and execution is exactly where most “bigger model” bets have historically stumbled.
It also says something about how the AI industry keeps score right now. A few years ago, the headline metric was benchmark scores. Then it became price per token. Now, at least for a moment, a handful of labs are back to competing on raw parameter count — the same metric the field spent 2024 and 2025 insisting didn’t matter as much as everyone thought. That’s not necessarily hypocrisy. It’s a sign that the labs still capable of funding a 10-trillion-parameter training run are running out of easier ways to differentiate, so they’re reaching for the one lever that’s expensive enough to be hard to copy.
So here’s the honest read: ByteDance training something this large is a real and notable signal of intent — it tells you the company is willing to spend enormous amounts of compute to be taken seriously at the frontier. What it doesn’t tell you, yet, is whether the model will actually be good. That verdict is three to six months of pre-training away, minimum, and the real test won’t be the parameter count on release day. It’ll be whether ByteDance can turn all that scale into something that actually outperforms models a fraction of its size — the same test every “biggest model ever” has had to pass, and not all of them have.
Leave a Reply