Home › Blog › The AGI Threshold
🧩

The AGI Threshold:
Are We Building AGI, or Just Getting Remarkably Close?

AGI is not a model that gets announced. It's a capability stack being filled in unevenly — sharp peaks, real valleys. The right question isn't which model wins. It's which pieces of the stack exist, and how fast the missing ones are closing.

FL
FrontierAGI Team
AGI Research Analysis

1. 🚀 What's Changed

Three years ago, the frontier conversation was about chatbots that occasionally hallucinated citations. Today it is about systems that plan multi-step tasks, call tools, write and execute code, browse and act on real interfaces, and — increasingly — keep working for hours without a human checking in. That pace of change is what has forced the AGI question back onto the table, not as science fiction, but as a live analytical question serious labs, investors, and policymakers are actually trying to answer.

0 → 9
Distinct capability layers now in active frontier development, up from essentially one (raw text prediction) a few years ago
Minutes → Days
Growth in unsupervised autonomous task duration across frontier agents in roughly two years
No Consensus
Definition of AGI or an accepted test for it, despite the accelerating capability curve
Jagged
The one word every serious researcher agrees describes today's frontier intelligence profile

This piece is deliberately not a model-versus-model comparison. It is an attempt to answer one question honestly: are we crossing into early AGI, or mistaking powerful agents for it? The answer, argued throughout, is that the question itself is slightly malformed — because AGI is not a single threshold you cross, but a stack of capabilities filled in at different rates.

2. 🧭 Defining AGI — The Undefinable Target

There is no accepted definition of AGI and no accepted test for it. This is not a minor academic gap — it is the central reason the debate stays unresolved even as capability visibly accelerates. The Turing Test is widely regarded as insufficient (modern chatbots pass casual versions of it without possessing anything resembling general intelligence). "Economically valuable work" definitions are more useful but circular — they define AGI by its effects rather than its nature. Benchmark-based definitions saturate faster than they can be agreed upon.

🎯 Narrow AI
Superhuman at one well-defined task (chess, protein folding, image classification) with zero transfer to unrelated domains. The default AI paradigm for most of the field's history, and still where most deployed systems sit.
🔀 General AI (Emerging)
Competent across a wide range of cognitive tasks without task-specific retraining — today's frontier LLMs sit here, partially, with capability that clearly transfers across domains but remains uneven and unreliable.
🧠 AGI (Contested)
Human-level competence across essentially all cognitively demanding tasks a human can do, with the reliability and adaptability to operate autonomously in the real world. No system has demonstrated this yet by most serious definitions — the contested question is how close current systems actually are.
⚡ Superintelligence (Speculative)
Intelligence substantially exceeding the best human performance across essentially every domain simultaneously. Purely theoretical at this stage — relevant to this piece mainly as the far end of the ladder, not as a near-term prediction.

Because no single definition commands consensus, this article adopts a working position: AGI is not one thing to detect, but a set of capabilities to measure separately. That reframing is the spine the rest of this piece is built on.

3. 📈 The Evolution: From Chatbots to Adaptive Autonomous Systems

The trajectory that makes the AGI question urgent now, rather than five years from now, runs through four fairly distinct stages — each defined by a genuinely new capability, not just a bigger model.

Conversational Stage
Reasoning Stage
Agentic Stage
Autonomous/Adaptive Stage
Conversational Stage 1
Chatbots
Fluent, single-turn or short-conversation text generation
CapabilityCoherent, contextually appropriate language generation across a huge range of topics — a genuinely new capability, but one confined to producing text in response to a prompt.
LimitationNo persistent goal, no tool use, no multi-step planning — every interaction starts fresh, and the model only ever tells you something, never does something.
Stage 1
Reasoning Stage 2
Reasoners
Extended deliberation before answering — test-time compute as a second scaling axis
CapabilityModels trained to think longer before responding, dramatically improving performance on math, code, and multi-step logic — the first sign that "more compute at inference" could buy genuine capability, not just speed.
LimitationReasoning remains largely confined to verifiable domains (math, code) where a clean reward signal exists — transfer to ambiguous, open-ended real-world judgment remains far weaker.
Stage 2
Agentic Stage 3
Tool-Using Agents
Intelligence plus action — calling tools, running code, browsing, checking results
CapabilityModels that don't just describe a solution but execute it — writing and running code, calling APIs, navigating software interfaces, and observing the result to decide the next step.
LimitationReliability degrades sharply as task length grows — a system that succeeds 95% of the time per step still fails most attempts at a 30-step task, the reliability problem examined in Section 9.
Stage 3
Autonomous/Adaptive Stage 4 — Now
Long-Horizon, Adaptive Systems
Pursuing goals across hours or days, adjusting strategy as conditions change
CapabilityThe current frontier: systems that maintain a goal across long stretches of unsupervised operation, adapting their plan as new information arrives rather than executing a fixed script.
LimitationThis is precisely the stage where the AGI question becomes unavoidable — and precisely where the capability stack in Section 4 is least filled in. Autonomy without reliability and self-awareness is not yet general intelligence; it's powerful, brittle agency.
Stage 4Current Frontier

4. 🧱 The AGI Capability Stack

If AGI isn't one thing, then no single benchmark can prove its arrival. What's needed instead is a layered framework — and a honest, current read of which layers are solid, which are partial, and which are barely started.

🛡️
Safe Autonomy — Knowing When Not to Act
Recognizing situations that require stopping, escalating, or asking for confirmation rather than proceeding.
Early
🪞
Metacognition — Calibrated Self-Awareness
Knowing what it doesn't know; accurately estimating its own confidence and error likelihood.
Early
🔁
Reliability & Self-Correction
Staying dependable across long chains of steps; catching and correcting its own errors before they compound.
Partial
📅
Long-Horizon Planning & Autonomy
Pursuing a goal across extended, unsupervised operation — minutes, hours, days — without losing coherence.
Partial
🖱️
Tool & Computer Use
Turning intelligence into action — running code, clicking through interfaces, checking real-world results.
Emerging-Strong
🧠
Memory & Continual Learning
Improving from lived experience rather than retraining from scratch; retrieval is not the same as learning.
Early
🌐
Multimodality → Coherent World Understanding
Text, image, audio, and video understood as one coherent world model, not four separately consumed input types.
Partial
🔀
Generalization
Transferring reasoning and knowledge into genuinely unfamiliar situations, not just recombining seen patterns.
Good-Partial
💡
Reasoning
Reliably solving novel, complex problems — reasoning when the path to the answer isn't already known.
Strong

Read from the bottom up, the pattern is unmistakable: the foundational, most-benchmarked layers (reasoning, generalization) are the strongest; the layers closest to real-world autonomous deployment (memory, metacognition, safe autonomy) are the weakest. That inversion — strong at the abstract, weak at the operational — is the single most important finding this stack surfaces, and it directly explains why demos can look like AGI while day-to-day deployment still requires heavy human oversight.

5. 🎙️ The Claims & Counterclaims

Frontier labs and their critics are, in a sense, both right — because they are usually talking about different layers of the same stack without saying so explicitly.

📢 The Optimist Case
Frontier labs' internal safety submissions and public roadmaps increasingly describe near-term systems as "broadly capable" across most cognitive tasks a knowledge worker performs
Reasoning and generalization — the two strongest layers in the stack above — are exactly the capabilities most associated historically with "general" intelligence
Autonomous task duration (Section 11) has grown at a pace that, extrapolated naively, reaches multi-week unsupervised operation within a few years
Economically, systems already substitute for large amounts of human cognitive labor across coding, analysis, and support work
🔍 The Skeptic Case
The weakest layers in the stack — memory, metacognition, safe autonomy — are precisely the ones required for genuinely independent operation, not the ones being showcased in demos
Jagged intelligence (Section 8) means systems ace competition-level math while failing tasks a child manages easily — inconsistent with any coherent notion of "general" competence
Extrapolating a capability curve naively has a poor historical track record — self-driving cars followed a similar early curve before hitting a long plateau of remaining edge cases
Economic substitution in narrow task categories is not the same evidence as general intelligence — a calculator substitutes for human arithmetic without being intelligent

Both sides are drawing on real evidence. The disagreement is less about facts and more about which layer of the stack should be given interpretive weight — reasoning and economic substitution (optimist emphasis), or reliability, memory, and metacognition (skeptic emphasis).

6. 🗂️ The Capability Matrix

Rather than ranking frontier models against each other, the matrix below maps which stack layers each demonstrates strong evidence for today — not who wins, but which pieces of AGI exist where. Ratings reflect independently observable capability patterns across the frontier model class as a whole, not vendor-specific benchmark claims.

Stack Layer Reasoning-Optimized Models Agentic/Tool-Use Models Multimodal-Native Models Long-Horizon Autonomous Systems
Reasoning
Generalization
Multimodal World Coherence
Memory / Continual Learning
Tool & Computer Use
Long-Horizon Planning
Reliability at Scale
Metacognition
Safe Autonomy

🟢 Strong  🟡 Good-Partial  🟠 Partial  🔴 Early/Weak — illustrative class-level pattern, not a single-vendor scorecard

The matrix's most important feature isn't any single cell — it's the visible column-wide weakness in the bottom three rows across every model category. No current class of frontier system has solved memory, metacognition, or safe autonomy well. That's the real finding, and it holds regardless of which specific model is fastest at any individual benchmark.

7. 📐 Why Benchmarks Can't Settle This

If AGI is not a single capability, then no single benchmark — however impressive the headline number — can prove its arrival. Three structural problems limit what any benchmark can tell us.

🎯 Benchmarks Are Narrow by Construction
A benchmark is, definitionally, a fixed, well-scoped task set. Excelling at it demonstrates competence at that task set — it cannot, by design, demonstrate the open-ended generality the AGI question is actually about.
📉 Benchmarks Saturate Fast
Frontier models now regularly saturate benchmarks within months of release, forcing constant replacement with harder tests — a moving target that says more about benchmark design velocity than about generality.
🕳️ Contamination Is Nearly Unavoidable
At internet-training-data scale, benchmark-adjacent content likely enters training corpora even unintentionally, inflating scores in ways that are extremely difficult to fully rule out.

None of this means benchmarks are useless — they remain the best available signal for tracking progress on individual stack layers. But no aggregate benchmark score, however high, should be read as direct evidence of AGI arrival. The gap between benchmark performance and the jagged, unreliable reality of open-ended deployment is the subject of the next section.

8. ⛰️ The Jagged Intelligence Problem

Incredible peaks, real valleys. The single most diagnostic pattern in frontier AI today is how spiky the capability profile actually is — superhuman on some tasks, subhuman on others that seem, intuitively, far easier.

Tasks Where Frontier Systems Are Superhuman
Competition mathematics
94
Code generation (scoped)
88
Information recall/synthesis
91
Formal logic puzzles
86
Tasks Where Frontier Systems Underperform a Child
Physical/spatial common sense
34
Robust novel-object manipulation
28
Consistent long-horizon judgment
31
Knowing when it's wrong
26

A four-year-old does not solve olympiad geometry — but reliably understands that a cup will fall if pushed off a table, recognizes when they don't know something, and doesn't confidently insist on a wrong answer with the same fluency they'd use for a right one. Frontier AI's profile is close to the mirror image of that. This isn't a minor quirk — it's evidence that current systems are not approximating human general intelligence along a single smooth axis. They are acquiring narrow superhuman peaks at a much faster rate than they are acquiring the broad, low-key robustness that underlies ordinary human competence.

The real question is not who's winning. It's which pieces of AGI exist, and how fast the missing ones are getting filled in.

9. 🔬 Gap Analysis: What's Actually Missing

🔁
Reliability at Long Horizons
A 99% per-step success rate collapses to under 40% survival probability over 100 sequential steps. This compounding-error dynamic, not raw capability, is the single largest practical barrier to autonomous deployment today.
🧠
Continual Learning
Retrieval and long context are not the same as learning. Humans update their models of the world from lived experience continuously; frontier systems remain largely frozen between training runs, however good their in-context adaptation looks over a single session.
🌍
Grounded World Models
Understanding physics, causality, and spatial relationships robustly enough to act reliably in the physical world remains unsolved — the exact gap the jagged-intelligence spatial/manipulation scores above illustrate.
🪞
Calibrated Self-Awareness
Knowing what it doesn't know — accurately, not just when explicitly prompted to hedge — remains the weakest layer in the entire stack, and arguably the one most necessary for trustworthy autonomous operation.

10. 🧭 How the Gaps Might Close

⏱️
Test-Time Compute Scaling
Addresses: Reasoning, Reliability
Letting a system deliberate longer and verify its own intermediate steps directly attacks compounding error — the second scaling axis behind reasoning-model progress.
💾
Persistent Memory Architectures
Addresses: Continual Learning
Systems that write durable, structured memory across sessions — not just longer context windows — are the most direct proposed path toward genuine experience-based learning.
🔌
Richer Tool Ecosystems
Addresses: Tool Use, Grounding
Standardized protocols connecting models to real software, sensors, and data sources widen what "acting in the world" can mean without requiring the core model itself to change.
🦾
Embodied & Multimodal Training
Addresses: World Models, Grounding
Training on physical interaction data, not just internet text and video, is the leading hypothesis for closing the spatial/causal reasoning gap the jagged-intelligence data shows most starkly.
🛡️
Safety & Oversight Frameworks
Addresses: Safe Autonomy, Metacognition
Treated here as a precondition for closing the autonomy gap, not an afterthought — a system that can't reliably recognize when to stop cannot be trusted with longer unsupervised horizons regardless of raw capability.
📊
Process Supervision at Scale
Addresses: Reliability, Metacognition
Training reward signals on the correctness of intermediate reasoning steps, not just final answers, is the emerging approach to teaching systems to catch their own errors before they compound.

11. ⏳ Autonomous Task Duration as the Honest Metric

Rather than tracking benchmark scores, the single most honest proxy metric for real progress may be simpler: how long can a system pursue a goal before a human needs to step in? Ten minutes, a day, a week — this trendline says more about genuine autonomy than any static leaderboard.

Seconds
2022
Single-turn
Minutes
2023
Multi-step chat
~1 Hour
2024
Early agents
Several Hours
2025
Coding/research agents
~1 Day
2026
Long-horizon agents
?
Next
Multi-day/week

Illustrative trend in unsupervised autonomous task duration before reliability degradation requires human correction

The trendline is genuinely steep — but extrapolating it naively into a firm AGI-arrival date is exactly the mistake this article is trying to avoid. Duration growth could continue smoothly, could hit a reliability wall (per Section 9's compounding-error math), or could advance unevenly across task types. It is a better metric than most, precisely because it directly measures the reliability and planning layers that matter most — not because it resolves the uncertainty.

12. 🌀 Recursive Self-Improvement, Examined Carefully

No concept in the AGI discourse is simultaneously more consequential and more overclaimed than recursive self-improvement — the idea of an AI system meaningfully improving its own capability without human-driven retraining, potentially compounding rapidly.

🧩 What It Would Actually Require
Reliable continual learning (Section 9's weakest layer) — a system that can't learn from experience can't meaningfully improve itself between training runs
Calibrated metacognition sufficient to correctly judge which self-modifications actually help versus merely appear to
Safe autonomy robust enough that self-modification doesn't require constant human verification, defeating the point
A feedback loop tight and reliable enough to compound meaningfully faster than ordinary human-driven research already does
📍 Where We Actually Are
AI-assisted research (models helping human researchers design experiments, write code, analyze results) is real and already happening today
This is meaningfully different from autonomous recursive self-improvement — a human remains firmly in the loop at every consequential step
The specific capabilities recursive improvement would require are, per the matrix in Section 6, among the weakest currently demonstrated
Most credible near-term progress looks like accelerated human-AI research collaboration, not autonomous compounding

The honest read: recursive self-improvement remains a coherent long-term concern worth taking seriously in safety planning, but current systems are far short of the specific capability combination it would require — and the gap is not merely quantitative (more compute), it is the same set of unsolved qualitative gaps (memory, metacognition, safe autonomy) examined throughout this piece.

13. 🔮 Three Futures, No Firm Timelines

Deliberately scenario-based rather than date-based — because the capability stack framework this article has built makes clear that a single confident date is not a defensible claim.

Scenario A
Early AGI, Now — Jagged but Real
The strongest stack layers (reasoning, generalization, tool use) continue improving fast enough that the weaker layers become manageable through better scaffolding and oversight, rather than needing to be independently solved. AGI arrives as an increasingly capable, increasingly trusted jagged system — never a clean single threshold, but broadly "AGI-equivalent" for most practical purposes within a few years.
Scenario B
AGI in Years — Gaps Close via Known Levers
The gap-to-trajectory levers in Section 10 (persistent memory, embodied training, process supervision, safety frameworks) each make steady, non-trivial progress over a multi-year horizon, gradually filling in the weak layers of the stack one at a time rather than all at once. AGI, if it arrives, does so unevenly across task domains — general in some areas years before others.
Scenario C
Fundamental Blockers Emerge
One or more weak layers — most plausibly continual learning or grounded world models — proves architecturally hard, not just data- or compute-hard, under current paradigms. Progress on reasoning and tool use continues, but genuine autonomy plateaus well short of anything resembling AGI, echoing how earlier "just scale it" predictions in other AI subfields eventually hit real walls.

Nothing in this article should be read as betting on any one of these three. The capability-stack lens is useful precisely because it lets you track which scenario the evidence is pointing toward, layer by layer, without needing to commit to a single narrative in advance.

14. 💼 Economic AGI vs. True AGI

For enterprises, the philosophical AGI debate is often beside the point. What matters practically is a narrower, more useful distinction: "economic AGI" — systems that are economically equivalent to general intelligence for a specific, bounded task class, regardless of whether they satisfy any deeper philosophical bar.

🧑‍💻
Copilots
Human drives, AI assists within a single task — the current dominant enterprise deployment pattern.
🤖
Agents
AI completes a bounded, well-scoped task end to end with human review at the boundaries, not every step.
🏢
Agentic Orgs
Multiple coordinating agents handle interlocking workflows across a function with light human oversight — the current frontier of enterprise deployment.
🌐
Possibly, AGI
The theoretical endpoint of this ladder — general enough to require no task-specific scoping at all. Where the philosophical debate this article has covered actually gets tested against reality.

Enterprises adopting this ladder don't need to resolve the AGI definitional debate to make good decisions — they need to correctly identify which rung a given task sits on today, and plan investment accordingly. The economic ladder advances independently of, and usually ahead of, the philosophical question of "true" AGI.

15. ⚖️ AGI as a Threshold, Not a Date

🎯 Our Take

The capability stack framework this article has built points to one clear conclusion: AGI is not a model that gets announced on a Tuesday. It is a set of layers being filled in at visibly different rates — reasoning and generalization moving fast, memory and metacognition and safe autonomy moving much slower — and the honest state of the evidence today is jagged, not general.

That jaggedness is not a temporary embarrassment on the way to a clean threshold. It may simply be what the arrival of AGI looks like from the inside — uneven, contested, visible in some domains years before others, filled with systems that are simultaneously superhuman and unreliable in ways that resist a single tidy verdict.

The same gradient logic extends naturally into physical and embodied AI — grounded world models and robust real-world manipulation sit even further behind the digital-cognitive layers mapped in this piece, suggesting that whatever threshold eventually gets crossed will be crossed unevenly there too, on its own delayed timeline.

For builders, enterprises, and policymakers, the practical implication is the same one this piece has returned to throughout: track the stack, not the headline. Watch which layers are closing and which remain stubbornly open, favor the economic-AGI ladder over the philosophical debate for near-term decisions, and treat any confident date-certain AGI prediction — in either direction — with real skepticism. The threshold is real. It just isn't a date.