Google is doing something unusual in a market obsessed with shipping faster: it tore down a frontier model and rebuilt it from scratch. Reporting indicates the company scrapped the base architecture of Gemini 3.5 Pro for a ground-up rebuild aimed at closing specific capability gaps — and, in the same stretch, quietly confirmed that Gemini 4 has entered pre-training. Together the moves reveal a lab willing to slow down to fix foundations even as it lines up its next flagship.

Why Rebuild a Frontier Model?

According to third-party reporting and leaks, Google concluded that Gemini 3.5 Pro's existing architecture could not close three stubborn gaps: mathematical reasoning, scalable vector graphics (SVG) scene generation, and overall image quality. Rather than patch around those limits, the company reportedly rebuilt the base model.

The rebuilt version is said to feature a 2-million-token context window — double the prior generation's one-million cap — along with a Deep Think reasoning layer for multi-step logic and autonomous workflow capabilities for chaining complex coding and tool-use tasks. These specifications come from leaks rather than official Google documentation, so they deserve caution until confirmed. But the strategic signal is clear: Google is prioritizing hard reasoning and agentic tool-use over incremental polish, betting those are the capabilities that will separate frontier models in the next phase.

The broader Gemini 3.5 family reflects the same thesis. Google has positioned its Flash tier as bridging the speed-versus-intelligence tradeoff — pairing frontier-level reasoning with fast, low-cost inference — and built it explicitly for complex agentic workflows and coding, the workloads now driving enterprise demand.

Reasoning Is the New Battleground

The rebuild lands amid a reshaped benchmark landscape. The simple leaderboards of two years ago have given way to a richer, harder set of tests designed to resist memorization:

  • Humanity's Last Exam (HLE) — probing the absolute frontier of knowledge.
  • GPQA Diamond — PhD-level science questions.
  • ARC-AGI-2 — novel reasoning that cannot be pattern-matched from training data.
  • FrontierMath Tier 4 — among the hardest mathematical problems available.

In this environment, models are judged less on fluent prose than on genuine multi-step reasoning and the ability to solve problems they have never seen. Google's decision to rebuild specifically around mathematical reasoning is a direct response to where the competitive bar has moved.

Gemini 4 on the Horizon

Buried in the launch of Gemini 3.6 Flash, Google confirmed on July 21, 2026 that Gemini 4 is in pre-training, describing it as the company's "most ambitious pre-training run yet." CEO Sundar Pichai has called it "significantly larger" than any prior Gemini, with coding and autonomous agents named as priorities.

What is known remains deliberately narrow: there is no release date, no published benchmarks, no pricing, and no context-window figure. The confirmation is less a product announcement than a marker — a statement that Google intends to keep pace at the very top of the frontier even as rivals ship in rapid succession.

Why It Matters

The Gemini rebuild cuts against the prevailing rhythm of 2026, in which labs have shipped new flagship models within days of one another and ignited a fierce price war. Choosing to rebuild a base model from scratch is expensive and slow — a bet that durable architectural foundations matter more than winning any single release cycle. For a company competing with OpenAI, Anthropic and a wave of open-weight Chinese models, it is a wager that quality of reasoning, not cadence of releases, will decide the next round.

For enterprises and developers, the practical implications are concrete. A frontier model rebuilt around mathematical reasoning, a 2-million-token context window, and agentic tool-use targets exactly the workloads — complex coding, long-document analysis, autonomous workflows — where reliability has been the binding constraint. If the rebuilt Gemini 3.5 Pro delivers on those gaps, it strengthens Google's hand in the enterprise AI contest that increasingly rewards models that can reason and act, not merely converse.

The dual disclosure also frames Google's roadmap as a relay: fix the current generation's foundations while the next, larger model trains in the background. Whether that patience pays off depends on execution neither the leaks nor the teaser can yet confirm. But in a field sprinting for headlines, a lab willing to rebuild rather than rush is itself a notable data point.

Sources