Home › Blog › The Gemini Series on the AGI Stack
🔵

The Gemini Series on the AGI Stack:
Google DeepMind's Long Game

The third lab in this series had the longest research pedigree and the latest product start. Tracing Bard/PaLM 2 through Gemini 3 shows a lab that entered the race behind on shipping, then used its research depth — multimodality, agentic browser prototypes, native tool use — to catch up on different levers than the other two labs used.

FL
FrontierAGI Team
Google DeepMind Model History AGI
How to read this post. Third in our per-lab lineage series — see The GPT Series and The Claude Series for the first two installments and the shared framework. Release dates and capability claims below are sourced from Google/Google DeepMind's own announcements and maintained release-history pages; lever/layer classifications are our editorial analysis.

1. 🧭 The Third Lab, a Different Starting Position

Google's position at the start of this race was almost the inverse of OpenAI's and Anthropic's: the deepest institutional research pedigree (DeepMind's AlphaGo/AlphaFold lineage, years of published transformer and RL research) paired with the slowest initial product response to ChatGPT. Bard's rocky February 2023 debut — a factual error in its own launch demo wiped billions off Alphabet's market cap — is a useful reminder that research depth and shipping speed are genuinely different capabilities, which is itself an early data point about which stack layers a lab is actually optimized for versus which ones it has to catch up on.

Feb 2023 → 2026
Bard's rocky debut to Gemini 3 — the shortest branded lineage of the three labs, but built on the longest research history
1M → 2M
Context window jump from Gemini 1.0 to Gemini 1.5 — Google's earliest and most distinctive lever
Dec 2024
Gemini 2.0 Flash Thinking arrives — roughly 3 months after OpenAI's o1 preview, Google's own test-time-compute entry
45–85%*
Gemini 3 Deep Think's disputed ARC-AGI-2 score range — see our Astra/Claude/Gemini post for the full discrepancy

2. 🧱 Quick Reference

Lever 1: Pre-training Scale Lever 2: Compute/Efficiency Lever 4: Post-Training Lever 3: Test-Time Compute Lever 5: Agentic Scaffolding Lever 6: Memory/Safe Autonomy Research

3. 📅 The Full Gemini Timeline

Pre-Gemini
Founding/Scale
Context/Efficiency
Agentic Era
Reasoning Era
Pre-GeminiFeb – May 2023
Bard (LaMDA → PaLM 2)
A rocky public debut, then a mid-year upgrade to a stronger base model
LeverBard launched (Feb 2023, public March 2023) on LaMDA, a weaker base than what OpenAI and Anthropic were already shipping. PaLM 2 (announced Google I/O, May 2023) was a straightforward Lever 1 scale-up used to catch the base model up to the field.
Lever 1
FoundingDec 6, 2023
Gemini 1.0 (Ultra, Pro, Nano)
The real founding release — three sizes, natively multimodal from the start
LeverGoogle's actual entry into the frontier race, a full year after ChatGPT. Unlike both other labs' founding releases, Gemini 1.0 was designed multimodal from the ground up (not bolted on later, as GPT-4V was) and launched in three tiers simultaneously — Google's version of Anthropic's day-one tiering bet, but centered on modality rather than just cost.
Lever 1
Context/EfficiencyFeb 2024
Gemini 1.5 (Pro, Flash)
1 million (later 2 million) token context window — the lab's defining early lever
LeverRather than chasing OpenAI/Anthropic's reasoning or agentic moves, Google leaned hardest on a different lever entirely: context length, backed by a Mixture-of-Experts architecture change for efficiency. This is Google's first genuinely distinctive lever choice in the series — nobody else was pushing context this far this early.
Lever 2Lever 1
Agentic EraDec 2024 – Feb 2025
Gemini 2.0 (Flash, Flash Thinking) + Project Mariner
Explicitly branded "the agentic era" — plus Google's own reasoning entry
LeverGemini 2.0 was framed by Google itself around agentic AI in multistage workflows. Gemini 2.0 Flash Thinking (announced Dec 19, 2024) arrived roughly three months after OpenAI's o1 preview — Google's test-time-compute entry. Project Mariner (research prototype, same window) explored an AI agent operating a real web browser — an early, distinctly Google approach to Lever 5 (agentic scaffolding via browser-level tool use rather than API function-calling).
Lever 5Lever 3
Reasoning EraMar – Jun 2025
Gemini 2.5 Pro / Deep Think
"Thinking" becomes standard, not experimental — reached GA June 17, 2025
LeverGoogle's own framing was explicit: Gemini 2.5 combined "a significantly enhanced base model with improved post-training" — a direct, publicly stated Lever 1 + Lever 4 combination claim, distinct from treating reasoning as purely a test-time trick. The experimental "Deep Think" mode (multiple hypotheses considered before responding) pushed further into Lever 3 territory.
Lever 1Lever 4Lever 3
Current Frontier2026
Gemini 3 (Pro, Deep Think, Flash) + Antigravity
Deployed everywhere at once — Search, the app, AI Studio, Vertex, a new dev platform
LeverGemini 3 Pro reached 1,501 Elo on LMArena; Gemini 3 Flash gained "agentic vision" (studying image details across multiple passes). The new Google Antigravity developer platform is Google's most explicit Lever 5 bet yet — leadership (Hassabis, Kavukcuoglu) frames the whole generation as a step toward AGI, echoing OpenAI's language more than Anthropic's.
Lever 5Lever 3Current Frontier

4. 🌱 Bard / PaLM 2: Catching the Base Model Up (2023)

Bard (LaMDA) → Bard (PaLM 2)
Feb 2023 → May 2023
Lever 1: Pre-training Scale
Bard's public launch was rocky by any account — a factual error in the announcement demo became a defining early story about the product, at a moment when both OpenAI's ChatGPT and Anthropic's Claude 1 were already shipping. The PaLM 2 upgrade three months later was a direct base-model scale-up specifically to close that gap — the clearest "catch-up" release examined anywhere in this series so far.
Stack layers touched: Reasoning, Generalization (both playing catch-up, not leading)

5. 🌟 Gemini 1.0: Multimodal From the Ground Up (Dec 2023)

Gemini 1.0 (Ultra, Pro, Nano)
December 6, 2023
Lever 1: Pre-training Scale
Google's real founding release, roughly a year behind ChatGPT and nine months behind Claude 1. The distinctive design choice: Gemini was natively multimodal from its first training run, rather than a text model with vision bolted on afterward (GPT-4's V variant arrived after GPT-4 itself; Gemini's multimodality was there at launch). Three tiers shipped together — Ultra, Pro, Nano — echoing Anthropic's day-one tiering instinct but organized around device/deployment target rather than purely price.
Stack layers touched: Multimodality (strong from day one, unusually early for the series), Reasoning

6. 📏 Gemini 1.5: The Context-Window Lever (Feb 2024)

Gemini 1.5 Pro / Flash
February 2024
Lever 2: Efficiency/ArchitectureLever 1: Scale
This is the release that establishes Google's most distinctive lever choice across the whole lineage: rather than chasing the reasoning or agentic moves the other two labs would make later, Gemini 1.5 pushed context length to 1 million tokens (later 2 million) — an order of magnitude beyond what any competitor was shipping at the time — backed by a Mixture-of-Experts architecture change for efficiency. Long-context "needle in a haystack" retrieval became Google's signature capability claim for most of 2024.
Stack layers touched: Memory (partial — long context is not the same as continual learning, but is a real practical substitute), Multimodality (video/audio understanding at scale)
While OpenAI and Anthropic were converging on reasoning and agentic scaffolding, Google spent 2024 on a lever nobody else was using seriously: just make the context window enormous.

7. 🤖 Gemini 2.0: "The Agentic Era" and Project Mariner (Dec 2024)

Gemini 2.0 Flash / Flash Thinking + Project Mariner
Dec 2024 – Feb 2025 (GA Feb 5, 2025)
Lever 5: Agentic ScaffoldingLever 3: Test-Time Compute
Google explicitly branded this generation around agentic AI in multistage workflows — a direct, public Lever 5 commitment. Gemini 2.0 Flash Thinking (announced Dec 19, 2024) was Google's own test-time-compute entry, arriving roughly three months after OpenAI's o1 preview (Sep 12, 2024) — a similar lag to Anthropic's five-month gap with Claude 3.7 Sonnet, reinforcing the pattern from our Claude post that this pivot was industry-timed. Project Mariner, a research prototype letting an agent operate a real web browser autonomously, was Google's own distinctive approach to tool use — browser-level control rather than API function-calling, foreshadowing the Antigravity platform.
Stack layers touched: Tool & Computer Use (browser-native approach, distinct from Claude's screen-reading computer use), Reasoning

8. 🧠 Gemini 2.5: Reasoning Becomes Standard (Mar–Jun 2025)

Gemini 2.5 Pro / Deep Think
March 2025 (GA June 17, 2025)
Lever 1Lever 4Lever 3
Google's own release framing is unusually explicit for this comparison: Gemini 2.5 combined "a significantly enhanced base model with improved post-training" — a direct, stated Lever 1 + Lever 4 combination, rather than crediting the gain to test-time compute alone the way the o1/3.7 Sonnet narratives did. The experimental "Deep Think" mode, which considers multiple hypotheses before responding, pushed further into Lever 3, foreshadowing the Deep Think branding that would carry into Gemini 3.
Stack layers touched: Reasoning (major), Generalization, Multimodality (long-context + native code generation)

9. 🔴 Gemini 3: Deployed Everywhere, AGI Language Included (2026)

Gemini 3 (Pro, Deep Think, Flash) + Antigravity
2026
Lever 5: Antigravity PlatformLever 3: Deep Think
Covered in detail in our Astra, Claude & Gemini post: Gemini 3 Pro reached 1,501 Elo on LMArena; Gemini 3 Flash gained "agentic vision" (studying images across multiple passes rather than one glance); Gemini 3 Deep Think's ARC-AGI-2 score is disputed between sources at 45.1% and 84.6% depending on tooling and verification method — the same kind of harness-dependent discrepancy that defined the GPT-6 Astra story. The new Google Antigravity developer platform is Google's most explicit Lever 5 investment yet, and DeepMind leadership (Hassabis, Kavukcuoglu) describe the generation using AGI-forward language closer to OpenAI's framing than Anthropic's more cautious "powerful AI" terminology.
Stack layers touched: Reasoning (disputed magnitude), Tool & Computer Use, Multimodality

10. 📊 The Full Picture: Gemini Releases × Stack Layers

ReleaseReasoningGeneraliz.MultimodalMemoryTool UseLong-HorizonReliabilityMetacog.Safe Auton.
Bard/PaLM 2
Gemini 1.0 native
Gemini 1.5 long ctx
Gemini 2.0 Mariner
Gemini 2.5
Gemini 3 disputed Antigravity undisclosed

🟢 Strong  🟡 Good-partial  🟠 Partial  🔴 Early/weak  ⚪ Not yet present — editorial synthesis based on sourced release claims above

The standout column is Multimodality — strong from Gemini 1.0 onward, earlier and more consistently than either the GPT or Claude matrices in this series. The weakest columns are the same two as every other lineage examined so far (Metacognition, Safe Autonomy) — but here they're additionally the two columns with the least public documentation of any lab in this series, per our earlier Astra/Claude/Gemini post's finding that Gemini's safety architecture disclosures are the thinnest of the three.

11. 🔁 Three Labs, Three Lever Sequences

Catch-Up Scale
Bard/PaLM 2
→
Native Multimodal Scale
Gemini 1.0
→
Context Window
Gemini 1.5
→
Browser-Native Agentic
2.0 + Mariner
→
Scale + Post-Training + TTC
Gemini 2.5
→
Platform-Wide Agentic
Gemini 3
🔵 Where Google Differs
Context length as a headline, standalone lever (Gemini 1.5) — neither OpenAI nor Anthropic made this their primary 2024 story the way Google did
Native multimodality from the founding release, earlier and more consistently strong than either other lab's matrix in this series
Browser-native agentic scaffolding (Project Mariner → Antigravity) as a distinct approach to Lever 5, versus Claude's screen-reading computer use and GPT's function-calling/tool APIs
AGI-forward public language from leadership (Hassabis, Kavukcuoglu), closer to OpenAI's framing than Anthropic's — despite disclosing the least about safety architecture of the three labs
🔶 Where All Three Converge
The test-time-compute pivot landed within a roughly 5-month window across all three labs (OpenAI Sep 2024, Google Dec 2024, Anthropic Feb 2025) — the strongest three-lab evidence yet that this transition was industry-timed, not any single lab's unique insight
Every current flagship (Astra, Opus 5/Fable 5.1, Gemini 3) leans hardest on Lever 3 + Lever 5 combined, not new pre-training scale
Metacognition and Safe Autonomy remain the two weakest columns across all three matrices in this series, with no lab past "good-partial" on either

12. 🔮 What's Next for Gemini

💾
Turning Long Context Into Real Memory
Google is arguably closest to a memory-lever head start of any lab in this series, given Gemini 1.5's context-length work — the open question is whether that becomes genuine continual learning or stays a (very large) retrieval window.
🌐
Antigravity as the Agentic Bet
Expect continued platform-level investment in Antigravity specifically, following the browser-native pattern set by Project Mariner rather than converging fully on Claude's or GPT's tool-use conventions.
🛡️
Safety Disclosure Is the Visible Gap
Given the ARC-AGI-2 discrepancy and the thinnest safety-architecture disclosure of the three labs, the most likely near-term correction is public pressure forcing more explicit safety-architecture statements, similar to what Anthropic already publishes routinely.

13. 🧭 Verdict

🎯 The Bottom Line
Google entered this race a year behind on shipping and caught up using a genuinely different lever than either competitor: not RLHF-style post-training refinement (OpenAI's InstructGPT moment) and not a front-loaded safety-post-training bet (Anthropic's Constitutional AI), but raw context-window scale plus native multimodality from its very first real release. That different starting lever produced a lineage that looks distinctive in exactly the columns you'd expect — Multimodality and, partially, Memory — and unremarkable in the ones every lab in this series still struggles with. Three labs, three different opening moves, and the same two unsolved layers at the top of the stack — which is itself the most important finding of running this comparison a third time: Metacognition and Safe Autonomy aren't lagging because any one lab under-invested. They're lagging industry-wide, regardless of starting position, founding philosophy, or which lever got pulled first.