1. 🧭 Three Different Bets on the Same Stack
The "Big Three" lineages (GPT, Claude, Gemini) share a starting resource: enormous private capital and enormous private compute. Grok, DeepSeek, and Qwen each started from a genuinely different resource constraint, which makes them a sharper test of the six-lever framework than another closed-lab comparison would be. xAI bet that raw compute scale, bought fast and applied aggressively, could compress years of catch-up into months. DeepSeek bet that architectural efficiency could substitute for the compute it didn't have unrestricted access to. Qwen bet that open-weight release velocity and breadth (many sizes, many languages, frequent updates) could win adoption even without leading every benchmark.
Nov 2023 → 2026Grok's fastest-compute-buildout lineage — Colossus 2's 1GW+ supercluster underwrites the entire strategy
37B / 671BDeepSeek V3's active vs. total parameters — the MoE efficiency ratio that shocked markets in Jan 2025
Aug 2023 → 2026Qwen's three-year, dozens-of-releases cadence — the fastest iteration speed of any lineage in this series
Huawei AscendDeepSeek V4's deliberate hardware target — a sovereign-compute lever no Western lab in this series has needed to pull
2. 🧱 Quick Reference
Lever 1: Pre-training Scale
Lever 2: Compute/Efficiency
Lever 4: Post-Training
Lever 3: Test-Time Compute
Lever 5: Agentic Scaffolding
Lever 6: Memory/Safe Autonomy Research
3. 🚀 The Grok Timeline: Compute First, Everything Else Later
Compute ScaleNov 4, 2023
Grok 1 (Beta)
Launched to X Premium users just months after xAI's founding
LeverPure Lever 1 — the fastest founding-to-first-release cycle of any lab in this entire series, made possible by starting with Musk's existing compute and capital access rather than years of prior research infrastructure.
Lever 1
Compute ScaleMar – Aug 2024
Grok-1.5 → Grok-2
128K context, then a full generational jump within five months
LeverGrok-1.5 (announced Mar 29, 2024) added a 128K context window and "improved reasoning" framing without a distinct reasoning-model architecture yet. Grok-2 (Aug 2024) was a straightforward scale-up, continuing the pattern of the fastest release cadence in the series relative to headcount and founding date.
Lever 1
ReasoningFeb – Jul 2025
Grok-3 (3T MoE) → Grok-4 (3T MoE)
A 3-trillion-parameter MoE architecture arrives, with reasoning framing
LeverGrok-3 (Feb 2025) and Grok-4 (Jul 2025) both used a 3-trillion-parameter Mixture-of-Experts architecture — Lever 1 scale combined with the same MoE efficiency architecture DeepSeek had already popularized, showing the efficiency lever's influence spreading across the industry even to a compute-first lab.
Lever 1Lever 2
Agentic/PlatformNov 2025 – Sep 2026
Grok 4.1 → Grok 4.6 → Grok 4.7
Iterating in place while Grok 5 trains on an expanding supercluster
LeverWith Grok 5 still training as of late 2026 on the Colossus 2 supercluster (1GW, expanding toward 1.5GW), xAI has iterated on the 4.x line (4.1, 4.6, 4.7) rather than pivot levers — a holding pattern while the next scale-up compute buildout completes, distinct from every other lab in this series, none of which has been this openly compute-gated on its next release.
Lever 1Lever 5
4. 🐋 The DeepSeek Timeline: The Efficiency Shock
Efficiency ArchitectureDec 26, 2024
DeepSeek V3
671B total parameters, only 37B active per token — the ratio that rattled markets weeks later
LeverThe clearest Lever 2 (compute/architecture efficiency) story in this entire series — Multi-head Latent Attention (MLA) plus a sparse MoE design let a 671B-parameter model activate only 37B per token, reportedly trained for a fraction of what comparable-capability Western models cost. This is the release that made "efficiency lever" a boardroom phrase, not just a research term.
Lever 2
RL/Post-TrainingJan 20, 2025
DeepSeek R1
RLVR-trained reasoning, open-weighted — the release that triggered the market selloff
LeverR1 popularized Reinforcement Learning from Verifiable Rewards (RLVR) at scale (building on Tülu 3's 2024 origin), reaching reasoning performance competitive with OpenAI's o1 at dramatically lower reported training cost, released open-weight under MIT license. This combination — Lever 3 (reasoning) plus Lever 4 (RL post-training) plus open access — is what actually moved markets in January 2025, more than any single benchmark number.
Lever 3Lever 4
RL/Post-TrainingMar – Sep 2025
DeepSeek V3-0324 → V3.1 Terminus
Rapid iteration refining the same efficient architecture rather than a new one
LeverContinued Lever 4 refinement on the V3 architecture — evidence DeepSeek treats post-training iteration as cheap and fast relative to a full architectural redesign, consistent with the efficiency-first philosophy set by V3 itself.
Lever 4
Sovereign ComputeNov 2025 – 2026
V3.2 / V3.2-Speciale / DeepSeekMath V2 → V4
A reasoning-first variant, then a deliberate pivot to non-Nvidia hardware
LeverV3.2-Speciale (reasoning-first) and DeepSeekMath V2 (Nov 27, 2025) continued the reasoning/post-training line. DeepSeek V4 (2026), covered in our Astra/Claude/Gemini post, is a roughly 1-trillion-parameter native multimodal MoE model deliberately optimized for Huawei Ascend chips rather than Nvidia — a lever no other lab in this series has needed to pull, driven by export-control realities rather than pure technical preference.
Lever 2Sovereign Compute
5. 🟠 The Qwen Timeline: Iteration Velocity as the Lever
Breadth/ScaleApr – Nov 2023
Tongyi Qianwen → Qwen-7B/14B/1.8B/72B
Public unveiling in April, then four size variants shipped within eight months
LeverAlibaba's founding lever wasn't a single flagship — it was breadth. Four distinct model sizes released within eight months of the initial public unveiling (April 11, 2023) established a "many sizes, fast cadence" pattern this lineage would repeat more consistently than any other lab in this series.
Lever 1
Breadth/ScaleFeb 2024 – Late 2024
Qwen 1.5 → Qwen 2
Eight size variants in one family release, then the model that shifted global attention to Chinese open weights
LeverQwen 1.5 (Feb 2024) shipped eight sizes (0.5B to 110B) simultaneously — an extreme version of the breadth lever. Qwen 2 (late 2024) is widely credited as the release that "shifted the community's attention toward Chinese-originated models" as serious frontier competitors, ahead of DeepSeek's more famous January 2025 moment.
Lever 1
MoE EfficiencyJan – Mid 2025
Qwen2.5-Max → Qwen 2.5 refinement
MoE efficiency refined, 100+ languages, active-parameter reduction
LeverQwen2.5-Max (Jan 29, 2025) and the broader 2.5 refinement pushed Lever 2 (MoE efficiency, drastically reduced active-parameter count) while expanding to over 100 languages — a breadth-plus-efficiency combination distinct from DeepSeek's efficiency-alone focus in the same window.
Lever 2
Frontier-Class PushApr 2025 – Aug 2026
Qwen 3 → Qwen3.5 (open) → Qwen3.8-Max
Bringing a genuinely Max-class model to open release for the first time
LeverQwen 3 (Apr 2025), Qwen3.5 (Feb 2026, open-weight alongside a proprietary Qwen3.5-Plus), and Qwen3.8 (Aug 2026, released Aug 3, 2026) — the team's own framing is that Qwen3.8 is the first time a genuinely Max-class model has been brought to open release, closing the gap this lineage had maintained between its open and closed-tier offerings.
Lever 1Lever 4
6. 🚀 Grok Deep Dive: The Compute-First Wager
xAI's Entire Strategy
Nov 2023 – Present
Lever 1: Compute Scale
No lab in this series has been this explicit about compute being the entire bottleneck. Grok 5's delay from Q1 to Q2 2026 was reported directly in terms of the Colossus 2 supercluster's buildout timeline (1GW, expanding to 1.5GW by April 2026), not a research or architecture obstacle. Elon Musk's own stated estimate — "Grok 5 will be AGI or something indistinguishable from AGI," with a self-reported 10% probability and rising — is worth reading literally: a 10% figure, stated by the person making the claim, implies a 90% chance it is not. This is the most compute-dependent, least architecture-differentiated lineage in the series — its lever choice really is just "more," applied faster than anyone else can physically build data centers.
Stack layers touched: Reasoning, Generalization (via scale) — with Metacognition and Safe Autonomy the least publicly documented of any lineage in this series
7. 🐋 DeepSeek Deep Dive: Efficiency as a Geopolitical Lever
V3 → R1 → V4
Dec 2024 – 2026
Lever 2: EfficiencySovereign Compute
DeepSeek's sequence inverts the usual order: rather than scale-then-efficiency (the pattern in every other lineage examined in this series), DeepSeek led with an aggressive efficiency architecture (V3's 37B/671B active/total MoE ratio) specifically because unrestricted compute access wasn't guaranteed — export controls made Lever 1 a constrained resource from day one, forcing Lever 2 innovation earlier and harder than any Western lab needed to. V4's explicit optimization for Huawei Ascend hardware over Nvidia extends this into a genuinely new lever this series hasn't needed to name before: sovereign hardware compatibility as a first-class design constraint, not an afterthought.
Stack layers touched: Reasoning (R1's RLVR breakthrough), Generalization, with efficiency touching every layer simultaneously
DeepSeek didn't just find a cheaper way to train a good model — it proved efficiency itself could be treated as a lever powerful enough to move markets, not just a cost-optimization footnote.
8. 🟠 Qwen Deep Dive: Winning on Breadth and Speed
The Full Qwen Family Strategy
Apr 2023 – Present
Lever 1: BreadthLever 2: Efficiency
Qwen's distinguishing lever isn't any single architectural innovation — it's release cadence and size breadth as a strategy in itself. Shipping four to eight size variants per major family release, in languages and use cases the "Big Three" often deprioritize, has made Qwen models a default choice for cost-sensitive and non-English deployment even when they don't lead individual benchmarks. Qwen3.8's explicit framing as "the first Max-class model brought to open release" signals the lineage catching up on Lever 1 (raw top-end capability) after building its reputation on Lever 2-adjacent breadth and efficiency.
Stack layers touched: Generalization (breadth across languages/domains), Reasoning (catching up via Qwen 3/3.8)
9. 📊 The Full Picture: Three Lineages × Stack Layers
| Lineage (Latest) | Reasoning | Generaliz. | Multimodal | Memory | Tool Use | Long-Horizon | Reliability | Metacog. | Safe Auton. |
| Grok 4.6/4.7 | | | | | | | | | least disclosed |
| DeepSeek V4 | | | native | | | | | | |
| Qwen3.8-Max | | breadth | | | | | | | |
🟢 Strong 🟡 Good-partial 🟠 Partial 🔴 Early/weak — editorial synthesis based on sourced release claims above
None of the three columns for Long-Horizon Autonomy, Metacognition, or Safe Autonomy get past "early/weak" for any of these three lineages — consistent with, but even more pronounced than, the pattern in the GPT/Claude/Gemini matrices. This tracks with a real resource explanation: safety architecture and long-horizon agentic scaffolding are exactly the layers that benefit most from sustained, well-funded, dedicated safety research teams — the resource the Big Three have invested in most visibly and that these three lineages, operating under compute constraints, export controls, or breadth-over-depth strategies, have had less occasion to prioritize.
10. 🔁 Six Labs, One Pattern
🆚 Where This Trio Differs From the Big Three
Compute itself is a named, public bottleneck for Grok in a way it never was for OpenAI, Anthropic, or Google — none of whom publicly blamed a release delay on data-center buildout timelines
DeepSeek and Qwen both treat Lever 2 (efficiency) as the primary lever, not a later optimization pass — the opposite ordering from every Big Three lineage, which all led with scale first
Sovereign hardware compatibility (DeepSeek V4's Huawei Ascend optimization) is a lever unique to this trio — none of the Big Three have needed to design around export controls
🤝 Where All Six Labs Converge
Every lineage examined across all four posts in this series shows the same two weakest columns: Metacognition and Safe Autonomy, regardless of funding model, compute access, or founding philosophy
MoE architecture has spread to every lineage in this series (GPT via unconfirmed reports, Claude undisclosed, Gemini's efficiency work, Grok's 3T MoE, DeepSeek's founding innovation, Qwen's refinement) — the closest thing to a universally adopted Lever 2 technique across all six labs
Reasoning-model/test-time-compute adoption has now reached all six lineages, though on very different timelines — DeepSeek R1 (Jan 2025) was actually faster to a competitive open reasoning model than either Grok or Qwen managed at the time
11. 🔮 What's Next for These Three
⚡
Grok 5 Is a Compute Story, Not a Research Story
Whether Grok 5 ships on time and delivers on Musk's stated (self-doubted) AGI framing depends almost entirely on Colossus 2's power buildout completing — the clearest example in this entire series of the power/grid constraint from our Six Levers post directly gating a release date.
🇨🇳
DeepSeek's Hardware Diversification Deepens
Expect continued investment in non-Nvidia-compatible architectures as export-control pressure persists — DeepSeek is likely to keep treating hardware portability as a permanent design constraint, not a one-off adaptation.
📈
Qwen Keeps Closing the Top-End Gap
Having established breadth and efficiency as its brand, expect Qwen's next releases to keep pushing top-end, Max-class capability specifically — following Qwen3.8's framing as evidence the lineage is actively closing, not just maintaining, the gap to frontier-tier reasoning scores.
12. 🧭 Verdict
🎯 The Bottom Line
Three lineages, three different starting resources, three different lever orderings — and the same two unsolved layers at the top of the stack as every lab examined in this series so far. xAI bet on compute scale applied faster than anyone else could physically build it, and is now visibly gated by the power/grid constraint this site has flagged as the meta-bottleneck underneath every lever. DeepSeek inverted the usual scale-then-efficiency order because export controls made scale a constrained resource from day one, and in doing so proved efficiency itself could function as a headline lever, not just a cost footnote — then extended that logic into sovereign hardware compatibility, a lever unique to this trio. Qwen won without leading a single benchmark category by treating breadth and iteration speed as the lever, then used that market position to fund a credible push at the top end. Six labs, six different combinations of the same six levers — and Metacognition and Safe Autonomy remain the weakest layer for every single one of them, which by this point in the series is no longer a lab-specific finding. It's the honest state of the entire field.