Google DeepMind reportedly scrapped the original Gemini 3.5 Pro base model and rebuilt it from scratch, according to multiple industry reports — not a routine tune-up, but a full restart after engineers found the model breaking down on core agentic tasks. The rebuilt version is now reportedly targeting a launch as soon as tomorrow, July 17, 2026, though Google itself has not confirmed a date.
The story, sourced from third-party reporting rather than any Google announcement, centers on why a frontier lab would throw away months of pre-training and start over — and what that says about how hard “agentic” AI actually is to get right.

What Went Wrong With the Original Model
According to reports from AIToolsRecap and TechTimes, Gemini 3.5 Pro’s original base model reportedly failed internal enterprise testing in three areas: mathematical reasoning, recursive tool-calling, and complex SVG scene generation. Two of those, in particular, were described as severe rather than cosmetic.
Recursive tool-calling — a model’s ability to call a tool, evaluate the result, and decide whether to call another tool based on that result, repeatedly, without losing track of the task — reportedly broke down under complexity. AIToolsRecap called this “the defining requirement for an agentic coding model,” meaning the failure wasn’t a minor bug so much as a disqualifying one for a model Google wants developers to trust with multi-step coding work.
The model also reportedly could not maintain structural consistency when generating complex, multi-layered SVG scenes — a benchmark that’s become an informal proxy for whether a model actually understands spatial and hierarchical structure, rather than pattern-matching flat outputs.
Why Google Chose to Start Over Instead of Patching It
Normally, a model that underperforms on specific benchmarks gets another round of post-training, fine-tuning, or reinforcement learning — cheaper and faster than starting from scratch. Google reportedly judged the failures too structural for that. If recursive tool-calling breaks down at the base-model level, no amount of fine-tuning fully fixes it; the limitation is baked into the architecture and training run itself.
That’s reportedly why Google chose a full pre-training restart instead — a decision one report described as “unusual” at frontier-model scale, because it’s expensive and slow by definition. Rebuilding a base model of this size reportedly cost hundreds of millions of dollars and cost Google months of delay; the model was originally expected around June 2026, according to earlier reporting, before slipping to its current target date.
The alternative — shipping a patched version with a known architectural ceiling — would have meant carrying that limitation into a flagship model set to be compared directly against OpenAI’s GPT-5.6 and Anthropic’s latest release. Google reportedly decided that risk wasn’t worth the time saved.
What the Rebuilt Gemini 3.5 Pro Reportedly Adds
Leaked details — again, not confirmed by Google — point to a substantially different model than the one that got scrapped. The most-repeated claim is a 2-million-token context window, which would be double the 1,048,576-token context already shipping in Gemini 3.5 Flash since May. If accurate, that would make it the largest production context window offered by any major lab, large enough in theory to hold an entire mid-size codebase or a year of support tickets in a single prompt.
Reports also describe a new “Deep Think” reasoning layer — a mode that reportedly trades response speed for extra inference-time computation on harder problems, aimed at multi-step planning and tool orchestration rather than quick answers. Some reporting places this behind a separate, more expensive tier.
The third recurring claim is autonomous multi-file coding: the model reportedly shifting from single-turn replies to longer workflows where it plans, executes, and refines a task across multiple files with less human oversight at each step — the same “agentic coding” capability that reportedly broke the original model in the first place.
What’s Still Unconfirmed
It’s worth being blunt about how much of this rests on leaks rather than official statements. Google has not confirmed the July 17 launch date, the 2-million-token context window, the Deep Think layer, the autonomous coding claims, or any pricing figures circulating alongside them. One report as recently as July 13 carried the headline framing that “every spec remains unconfirmed.” Reporting on the leaks themselves isn’t even unanimous — one leak thread claims the rebuilt model beats rival models on coding and reasoning, while a separately sourced leak claims it still lags behind. Treat all of it, including the date, as reported and not settled until Google publishes a model card or an official post.
Why This Matters for Google vs OpenAI vs Anthropic
Scrapping a frontier pre-training run is rare enough that the decision itself is a data point about how competitive the current model race has become. Google is reportedly racing to answer OpenAI’s GPT-5.6 family and Anthropic’s newest flagship on the same battleground — agentic coding and tool use — rather than on chat quality alone, with DeepSeek’s own July 24 release adding pressure from a different direction entirely.
If you want to see how the field stacks up so far, Anthropic’s own Claude Sonnet 5 has already staked its claim on developer-focused agentic coding, and xAI’s Grok 4.5 is positioning itself as an Opus-class competitor on the same turf. We’ve also broken down OpenAI’s GPT-5.6 family in detail, since it’s the model Google is reportedly measuring itself against most directly.
None of that makes the rebuild story any less notable on its own. A lab throwing away a finished frontier model over agentic failures — rather than shipping it with a caveat — is as much a signal about where the industry thinks the real competition now lives as anything in the leaked spec sheet.
Deixe um comentário