Home
›
Blog
›
The AGI Debate: Scaling vs. New Architectures
Standalone Deep Dive · Fourth in the Evolution / After Transformers Arc
⚖️
The AGI Debate: Scaling vs. New Architectures — Steelmanning Both Sides
"After Transformers" named this as the single most consequential live disagreement in the field and moved on. This article stops and sits in it: the strongest honest version of the scaling camp's case, the strongest honest version of the new-architectures camp's case, where they actually agree, and what it would take to prove either one right — without declaring a winner that the evidence doesn't yet support.
FL
FrontierAGI Team
September 12, 2026 · 65 min read
Why This Deserves Its Own Article Instead of a Paragraph
"After Transformers" listed world models as an open bet and noted, in passing, that it's "a minority position relative to 'scale the current recipe further,' championed seriously by a researcher with deep credibility." That sentence compresses a genuinely important, genuinely unresolved argument into a single line — and compressing it further would do it a disservice. This article expands it properly: not to pick a side, but to show what a well-argued disagreement between serious researchers actually looks like, since you'll need to navigate exactly this kind of disagreement constantly as a researcher yourself.
A ground rule before starting: both positions below are held by credentialed, serious researchers with real evidence behind them, not a "mainstream view vs. a crank." Presenting one side as obviously correct would misrepresent the actual state of the field.
The existence of this debate, among researchers this capable, holding this much relevant evidence, is itself the strongest available evidence that the honest answer right now is "genuinely unresolved" — not a dodge, but the actual state of the field.
Part 1 — The Scaling Camp
The Case for "The Current Recipe, Scaled Further, Is Enough"
📈The Scaling Camp's Core Claim
The decoder-only Transformer, trained on more data with more parameters and more compute, augmented with practical patches (RAG, Mixture-of-Experts, test-time reasoning compute — all covered in "After Transformers") when a specific limitation shows up, is sufficient to reach dramatically more general capability. No fundamentally new architecture is required — the same recipe that produced the entire chain from GPT-1 to today's reasoning models keeps paying off when pushed further.
The Strongest Evidence For It
The empirical track record, traced across this site's own Evolution article: every specific limitation identified along the way (GPT-3's unreliable instruction-following, the knowledge-cutoff problem, weak multi-step reasoning) was solved by an addition to the existing recipe — RLHF, RAG, test-time compute — rather than a wholesale architectural replacement. Richard Sutton's widely cited 2019 essay,
"The Bitter Lesson", makes the sharpest version of this argument: across AI's history, general methods that leverage more computation have repeatedly beaten more clever, human-knowledge-encoded methods, and researchers attempting to build in domain-specific structure have repeatedly been embarrassed by simpler, more scalable approaches catching up and surpassing them within a few years.
Who's Betting Hardest Here, and Why
Labs and researchers with the largest compute budgets and the strongest empirical track record from the current recipe have the most direct evidence for continuing it, and the infrastructure economics reward a strategy that's already working over an unproven architectural bet — a rational, not merely conservative, position given the evidence in hand.
What Would Prove This Camp Right
Continued, unbroken progress on the specific capabilities the new-architecture camp claims are structurally out of reach (Part 2 below) — genuine continual learning, robust causal reasoning, dramatic sample-efficiency improvements — achieved through scaling and modest additions alone, without requiring a fundamentally different mechanism.
Part 2 — The New-Architectures Camp
The Case for "A New Mechanism Is Required"
🧩The New-Architecture Camp's Core Claim
Predicting the next token in text is a fundamentally limited proxy for genuine understanding of how the world works — it optimizes for statistical plausibility of text, not for an internal, causal model of reality. This camp argues certain capabilities (genuine continual learning without catastrophic forgetting, robust reasoning about physical cause and effect, human-like sample efficiency) are structurally out of reach for the current recipe, however far scaled, and require a different learning paradigm.
The Strongest Evidence For It
Yann LeCun's
"A Path Towards Autonomous Machine Intelligence" (2022) makes this case in detail: he points to the enormous gap in sample efficiency between how much data a human child needs to learn basic physical intuitions versus how much text data current models require to approximate similar competence, arguing this gap is evidence of a missing mechanism (a predictive "world model"), not just a missing quantity of data. Separately, catastrophic forgetting ("After Transformers," Part 3) remains genuinely unsolved by scaling alone — bigger models don't straightforwardly retain new information better without overwriting old information, which is exactly the kind of structural limitation this camp points to.
Who's Betting Hardest Here, and Why
Researchers whose core thesis is that generalization, not scale, is the actual bottleneck — this site's own
SSI investigation documents Ilya Sutskever's stated position that pretraining-as-usual is running into diminishing returns and that a different research direction is needed, a notable case of a researcher who was centrally involved in the scaling camp's own historical successes (Evolution article, Era 12) now betting elsewhere. This is a stronger form of evidence than an outsider's critique — it's a credible insider's own reassessment.
What Would Prove This Camp Right
A demonstrated capability — genuine continual learning, or dramatically more sample-efficient learning of physical/causal reasoning — achieved specifically through a new architectural mechanism, that scaling the existing recipe further demonstrably could not reach even with substantially more compute and data.
Part 3 — Where They Actually Agree
The Overlap Both Camps Would Accept
Debates are usually presented as pure disagreement, which overstates how far apart the two camps actually are. Several things both sides accept:
✅ Data quality matters, not just quantity
Chinchilla's finding (Article 1, Article 3) that data-to-parameter ratio matters is accepted by both camps — the scaling camp treats it as a refinement within the existing recipe; the new-architecture camp treats it as early evidence that raw scale alone was never the full story.
✅ Test-time compute and reasoning are a genuine, valuable axis
Reasoning models ("After Transformers," Part 1) have been adopted rapidly and are not controversial between the two camps — both treat spending more compute at inference time on harder problems as clearly useful, whatever their disagreement about pretraining's ceiling.
✅ Catastrophic forgetting and true continual learning remain unsolved
Neither camp claims this is solved today — the disagreement is specifically about whether scaling the current recipe will eventually solve it, or whether it structurally cannot.
✅ Hybrid, not purely either-or, approaches are worth pursuing
Most frontier labs, in practice, pursue scaling as the primary strategy while also funding some architectural and world-model research — the public "debate" framing is sharper than most labs' actual, hedged research portfolios.
Part 4 — Applying the Framework to the Debate Itself
Running "How to Read a Paper Like a Researcher"'s Framework on the Debate
This debate is itself a good subject for the six-question framework from "After Transformers" and "How to Read a Paper Like a Researcher" — not applied to one paper, but to the disagreement as a whole.
Is this an empirical disagreement or a definitional one?
Genuinely mixed — some of it is empirical (will scaling eventually solve catastrophic forgetting? — a question actual experiments could resolve), and some of it is closer to definitional (what counts as "genuine understanding" versus "sufficiently good statistical approximation" is a question reasonable people can disagree about even given identical evidence).
What would falsify each camp's position?
The scaling camp's position is falsified by a capability that scaling demonstrably cannot reach even with dramatically more compute; the new-architecture camp's position is falsified by scaling eventually reaching that same capability without a new mechanism. Notably, both falsification conditions require evidence that doesn't fully exist yet — which is exactly why the debate remains open rather than settled.
Is either side's argument load-bearing or motivated by incentives?
Worth naming honestly: compute-rich organizations have a real incentive to believe scaling continues to work, and researchers pursuing new architectures have an incentive to believe the current paradigm has a ceiling. Neither incentive, by itself, makes either position wrong — but it's a reason to weight the underlying evidence (Parts 1–2 above) more heavily than either camp's confidence in its own position.
Part 5 — If You're Choosing Where to Place Your Own Bet
What This Means for a Researcher Deciding Where to Work
This isn't a purely academic question if you're early in a research career and choosing a direction — per this site's Research Frontier Map and lane framework, which camp a lab or team is betting on shapes what kind of work you'd actually be doing day to day.
1Betting on scaling means working on infrastructure, data pipelines, and the practical add-ons (RAG, MoE, reasoning) covered in "After Transformers" — closer to this site's Lane 1 (frontier-scale) and Lane 2 (applied/agentic) profiles, with a faster empirical feedback loop since the recipe is proven.
2Betting on new architectures means working on genuinely open, higher-risk research — closer to Lane 3 (narrow research bet), with a slower feedback loop, higher uncertainty, and potentially higher long-term payoff if the bet is correct.
3Neither choice is obviously correct — and per Part 3 above, most real research careers involve some mix rather than a pure commitment to one camp, especially early on while you're still building the foundational skills covered in this site's Technical Stack and Self-Assessment articles.
A researcher who can articulate both camps' strongest arguments, and name specifically what evidence would change their own mind, is already doing better research judgment than one who has simply picked a side and stopped looking at the other camp's evidence.
⚠️ What's Uncertain in This Article
This article does not, and should not, declare a winner. As of September 2026, this is a live, contested disagreement among serious researchers, not a settled question this site is positioned to resolve. The framing of "two camps" is also a simplification — many researchers hold more nuanced, hybrid positions than either pure summary above fully captures, and this article's own characterization of specific researchers' views (LeCun, Sutskever) is a good-faith summary of their public statements and papers, not a verified restatement approved by them.
🎥 Recommended Videos
🧭 Closing — Holding a Genuine Disagreement Without Resolving It Prematurely
🎯 The Bottom Line
The scaling camp has an unbroken empirical track record and a sharp philosophical argument (the Bitter Lesson) behind it; the new-architecture camp has a credible insider's reassessment (Sutskever) and a specific, named structural gap (sample efficiency, continual learning) that scaling hasn't yet closed. Both positions survive the six-question framework's scrutiny — which is precisely what makes this a genuine, unresolved debate rather than a settled question with one side simply not paying attention. The skill this article is really asking you to practice isn't picking a side; it's the harder, rarer skill of holding a real disagreement clearly, naming what would resolve it, and updating honestly as new evidence — some of which you may end up producing yourself — arrives.