Home
โบ
Blog
โบ
AGI Benchmarks: How the Field Tries to Measure the Thing It Can't Yet Define
Reference Series ยท Sequel to the Benchmark Landscape
๐งญ
AGI Benchmarks: How the Field Tries to Measure the Thing It Can't Yet Define
This site's Benchmark Landscape article mapped capability testing โ reasoning, coding, agentic tool-use, and more. This sequel covers a narrower and stranger problem: the handful of benchmarks and frameworks explicitly built to measure progress toward AGI itself, despite the field having no agreed definition of what AGI is. It covers Humanity's Last Exam, the ARC-AGI family through its 2026 shift to interactive game environments, METR's autonomy research as a de facto operational AGI proxy, the Turing Test's afterlife, and the growing academic literature that critiques the entire benchmarking ecosystem on Goodhart's Law grounds.
FL
FrontierAGI Team
September 18, 2026 ยท 55 min read
You Can't Benchmark What You Haven't Defined
Every capability benchmark in this project's earlier Benchmark Landscape article measures something concrete: can this model solve a competition math problem, patch a real GitHub issue, click the right pixel on a screenshot. AGI benchmarks are a different animal entirely โ they are attempts to measure progress toward a destination that the field has never agreed how to define. OpenAI defines AGI in economic terms. Google DeepMind defines it as a six-level spectrum of performance and generality. Anthropic largely avoids the term. A 2025 paper with 33 co-authors proposed yet another definition grounded in human cognitive-science theory. Every AGI benchmark covered in this article is, underneath its methodology, a bet on which of these competing definitions matters most โ and that is worth understanding before trusting any single score.
The Definition Problem
Four Labs, Four Definitions
OpenAI โ Economic Value
OpenAI's charter defines AGI as "highly autonomous systems that outperform humans at most economically valuable work." This economic framing was not just philosophical โ it was contractually load-bearing: OpenAI's original 2019 partnership with Microsoft tied Microsoft's IP and commercial access limits to an OpenAI board declaration of AGI under this exact definition. In October 2025, the two companies signed a new agreement replacing a unilateral board declaration with verification by an independent expert panel, and extending Microsoft's IP rights through 2032 including post-AGI models under agreed safety guardrails.
Google DeepMind โ Levels of AGI
A 2023 paper by Shane Legg and collaborators ("Levels of AGI for Operationalizing Progress on the Path to AGI") proposes two independent dimensions โ Performance (how deep the skill is relative to humans) and Generality (how broad the skill is across tasks) โ with six performance levels: No AI, Emerging, Competent, Expert, Virtuoso, and Superhuman. Under this framework, AlphaFold and Stockfish sit at Superhuman but Narrow; general chat assistants of the ChatGPT/Bard/Llama 2 era were placed at only Emerging but General โ illustrating that "how good" and "how broad" are genuinely separate axes, not one scale.
Anthropic โ Deliberately Avoiding the Term
Dario Amodei's essay "Machines of Loving Grace" explicitly prefers "powerful AI" to "AGI," defining the threshold as intellectual capability exceeding that of Nobel laureates, genuinely multimodal, and able to act autonomously in the world rather than merely answer questions โ a deliberate rhetorical choice to avoid a term Anthropic considers under-specified and overloaded with science-fiction connotation.
The Tegmark Definition (2025)
A October 2025 paper with 33 co-authors โ including Dan Hendrycks, Yoshua Bengio, Gary Marcus, Eric Schmidt, and Max Tegmark โ proposed defining AGI as matching the cognitive versatility and proficiency of a well-educated adult, grounded explicitly in the Cattell-Horn-Carroll theory of human cognitive abilities from psychometrics โ an attempt to borrow a more rigorous, pre-existing measurement framework from human intelligence research rather than inventing a new one from scratch.
Why This Matters for Every Benchmark Below
A benchmark built around OpenAI's economic-value definition (can this system do valuable work) looks completely different from one built around Chollet's skill-acquisition-efficiency definition (can this system learn new things quickly from little data) โ and both differ from METR's autonomy-duration framing (how long can this system work unsupervised). None of the benchmarks in this article claims to have solved AGI measurement; each has picked one facet of a contested, multi-part problem.
Humanity's Last Exam
The Last Exam Before Knowledge Testing Runs Out
Humanity's Last Exam (HLE)Jan 2025
- Origin
- Created by the Center for AI Safety (CAIS) and Scale AI, released January 2025 as a deliberately extreme closed-form academic exam โ 2,500 expert-written questions spanning 100+ disciplines, explicitly designed to be the last broad-knowledge benchmark the field needs before evaluation must shift toward agentic and real-world tasks.
- Launch Scores
- Extremely low at release โ GPT-4o scored 2.7% and Claude 3.5 Sonnet scored 4.1%, versus a claimed expert human baseline around 90%, illustrating just how hard the question set was calibrated to be.
- Known Quality Issues
- Independent review by FutureHouse found roughly 29% of chemistry and biology questions may have incorrect reference answers; Scale AI's own estimate puts the overall error rate around 18%. A follow-up verification effort, "HLE-Verified," has attempted systematic review and correction of the original question set.
HLE's own creators are explicit that a perfect score would not itself indicate AGI โ it measures structured, closed-form academic knowledge recall and reasoning, not autonomy, creativity, or real-world task competence. This self-imposed caveat is unusually candid for a benchmark and is worth taking at face value: HLE answers "how much expert knowledge does this model have," not "is this AGI."
A Note on Current Scores
Reported figures for models tested in early-to-mid 2026 (for example, one secondary source citing Gemini 3 Pro Preview around 37.5% and GPT-5 Pro around 31.6% as of February 2026) come from aggregator trackers rather than a primary HLE leaderboard screenshot, and should be treated as directionally informative rather than exact. Several other figures found during research for this article were attributed to model names that could not be corroborated against any confirmed lab release and are excluded here entirely rather than repeated.
The ARC-AGI Family
From Static Puzzles to Interactive Games
Of every benchmark in this article, ARC-AGI is the one most explicitly built around a specific theory of intelligence. Franรงois Chollet's 2019 paper "On the Measure of Intelligence" argues that intelligence should be measured as skill-acquisition efficiency relative to prior experience โ not the possession of any particular skill, however impressive, but the rate at which a system can acquire a new skill from limited exposure. ARC-AGI's puzzle-grid tasks were designed so that each one is trivial for a human seeing it for the first time, yet resistant to the kind of large-scale pattern memorization that lets language models ace ordinary knowledge benchmarks.
2019ARC-AGI-1 (Chollet)
2025ARC-AGI-2, a harder successor
Mar 25, 2026ARC-AGI-3: interactive game environments
2025ARC-AGI-2
Introduced specifically because ARC-AGI-1 was approaching saturation for the strongest reasoning models. The gap was stark at launch: a leading OpenAI reasoning model that scored 75.7% on ARC-AGI-1 managed only about 4% on ARC-AGI-2 โ confirming the new version was a genuinely harder, not just relabeled, successor. The 2025 ARC Prize competition ran with $1 million in prizes across 1,455 competing teams; the best score on the private evaluation set reached only 24%, and no entrant reached the competition's target threshold.
Mar 25, 2026ARC-AGI-3
A genuine design shift, not just a harder puzzle set: hundreds of hand-crafted, turn-based, interactive game-like environments with no stated instructions or goals, requiring a model to explore, plan, remember, and discover its own objective โ capabilities static single-shot puzzles cannot test at all. At launch, frontier AI systems scored under 1% while human testers solved essentially all of the environments, the largest human-AI gap this project has seen reported on any benchmark. The 2026 ARC Prize competition carries more than $2 million in prizes, roughly double the prior year's pool.
Every other benchmark in this article's parent piece asks "can you produce the right output." ARC-AGI-3 asks "can you figure out what the right output even is" โ arguably a closer approximation of what general intelligence actually requires.
METR's Time Horizon as Proxy
"AGI" Redefined as a Duration, Not a Score
This site's Benchmark Landscape article already covered METR's headline finding โ the length of task an AI agent can complete autonomously at 50% reliability has roughly doubled every seven months since 2019, with evidence of acceleration to roughly four months more recently. What makes this research distinctly AGI-relevant, rather than just another capability metric, is METR's own explicit framing of a "one-month AGI" threshold: a system that can outperform humans given roughly 167 working hours โ about one month of full-time work โ chosen specifically because that is roughly the period a new human hire needs to onboard into a role and begin producing independent value. A system clearing that bar would, by this framing, be matching a genuinely human capacity for sustained, high-context, unsupervised contribution โ an operational definition of general capability that sidesteps the exam-question approach entirely in favor of a duration.
A Framing, Not a Consensus
This "duration as AGI proxy" framing is most clearly articulated in METR's own research writing and in adjacent AI-safety research communities. This article did not find a confirmed instance of a policymaker publicly adopting this exact framing as an official definition โ it should be understood as an influential research framing gaining traction, not a settled cross-institutional consensus.
RE-Bench & MLE-bench
Can the AI Do AI Research?
A distinct and safety-relevant category tests whether AI systems can perform the actual work of machine learning research and engineering โ directly bearing on recursive self-improvement discourse, since a system that can meaningfully accelerate its own successor's development crosses a threshold most AGI definitions treat as significant regardless of how it performs on exam-style questions.
RE-Bench (METR)Nov 2024
- Structure
- Seven open-ended machine learning research-engineering environments, benchmarked against a human baseline of 71 separate eight-hour attempts by 61 experts drawn from labs including DeepMind, Anthropic, and OpenAI, plus graduate students at Berkeley, CMU, Stanford, and MIT.
- The Striking Result
- When both AI agents (Claude 3.5 Sonnet and o1-preview at the time) and human experts were given the same short two-hour budget, the AI agents scored roughly four times higher than the human experts โ though this compares short-budget AI performance against humans using only a fraction of their own eight-hour allowance, not a full apples-to-apples long-horizon comparison.
MLE-bench, published by OpenAI, takes a complementary approach: 75 real Kaggle machine-learning competitions with established human leaderboard baselines. The best-performing setup found (an o1-preview model paired with an agent scaffold called AIDE) achieved a Kaggle bronze-medal-equivalent result in roughly 17% of the competitions tested โ genuine competence, but still a minority of cases, not the dominant outcome.
Composite Indices
One Score to Rule Them All?
Epoch AI's Capabilities Index attempts to solve the fragmentation problem directly: it combines scores across more than 50 individual benchmarks onto a single difficulty-anchored scale, using a statistical approach related to item-response theory, where models tested across many benchmarks help establish each benchmark's relative difficulty against the others. Epoch reports a strong internal fit (Rยฒ around 0.86โ0.87) for the resulting composite scale.
Why Composite Scores Remain Contested
Any single index inherits every weakness of its component benchmarks โ if a contributing benchmark is contaminated, saturated, or gamed, the composite score is too, just with the problem hidden one layer down. Epoch's own documentation acknowledges uncertainty in extending the index backward to pre-2023 models and in quantifying confidence at small sample sizes. No composite index has achieved the kind of cross-lab standard-setting status SWE-bench Verified achieved for coding specifically โ the field still largely reports individual benchmark scores rather than a single blended number.
The Turing Test's Afterlife
Passed, and Mostly Retired Anyway
A pre-registered, randomized controlled study from UC San Diego tested a modern three-party Turing Test setup: five-minute conversations, with a human judge trying to distinguish an AI from a real human interlocutor. With a persona-style system prompt, GPT-4.5 was judged to be the human party 73% of the time โ exceeding chance, and even exceeding the rate at which the actual human confederate was correctly identified as human. Without persona prompting, the same model's rate dropped sharply to 36%, and GPT-4o (no persona) scored only 21% โ showing the result depends heavily on prompting technique, not a fixed underlying capability.
Why the Field Mostly Shrugged
Despite this being a genuinely rigorous, pre-registered academic study โ not an informal demo โ the broader AI research community's reaction was muted. The consensus view is that the Turing Test measures conversational mimicry under adversarial persona-steering, not cross-domain reasoning, real-world task competence, or anything resembling the "economically valuable work" or "skill-acquisition efficiency" framings that dominate current AGI discourse. It is now widely treated as a historically important but practically obsolete AGI proxy โ gameable through prompting technique rather than reflecting the underlying capability differences that actually matter.
Survey Research on Benchmarking
The Papers That Critique the Whole Ecosystem
Beyond individual benchmarks, a growing body of academic meta-research examines whether the AI benchmarking ecosystem itself is trustworthy. Several papers explicitly invoke Goodhart's Law โ the principle that once a measure becomes a target, it stops being a good measure โ as the central organizing critique of how labs report and compete on benchmark scores.
"Can We Trust AI Benchmarks?"
An interdisciplinary review examining current issues in AI evaluation broadly, explicitly framing benchmark-chasing behavior through a Goodhart's Law lens.
Contamination as "the Dirty Secret"
A paper on standardizing AI evaluation from models to agents documents a striking contamination case: a major model reportedly memorized a benchmark's own anti-training "canary" marker โ text specifically inserted to detect whether a benchmark had leaked into training data โ direct evidence the safeguard itself had been ingested.
Systematic Evaluation-Safety Reviews
A PRISMA-guided systematic survey covering evaluation-safety failures in LLM benchmarking from 2018 through 2026 catalogs recurring methodological failure patterns across the field's history, treating benchmark unreliability as a chronic condition rather than isolated incidents.
Stanford HAI's AI Index
Functions as the field's closest thing to an institutional meta-survey, aggregating results across dozens of benchmarks annually. Its 2026 edition highlighted a striking illustration of the narrow-vs-general capability gap this article is built around: SWE-bench Verified performance rose from roughly 60% to near 100% within a single year, while a model capable of gold-medal-level International Mathematical Olympiad performance correctly read an analog clock only about half the time.
Why No Benchmark Has Won
The Fragmentation Is Structural, Not Temporary
This project's Benchmark Landscape article identified saturation as the default lifecycle of every capability benchmark โ introduced, adopted, saturated, replaced. AGI benchmarks show the same pattern at a higher level of abstraction: each new attempt (HLE, ARC-AGI-2, ARC-AGI-3) exists specifically because its predecessor stopped discriminating between top models, and each new attempt will very likely face the same fate eventually. The deeper reason is not a temporary measurement gap that better engineering will close โ it is that "AGI" itself has never been reduced to a single operational definition the field agrees on, so no benchmark built around any one definition can claim to have captured the whole concept. A benchmark built around Chollet's skill-acquisition-efficiency theory (ARC-AGI) will always look incomplete to someone who thinks AGI is fundamentally about economic value (OpenAI's definition), and vice versa.
What a Good AGI Benchmark Needs
Four Recurring Criteria From the 2025โ2026 Literature
1Contamination resistance โ genuinely novel, held-out, or continuously refreshed tasks, since any static public benchmark eventually leaks into training data.
2Economic or real-world grounding โ a score should connect to deployable value, not just abstract test performance, echoing OpenAI's own definitional framing even when the benchmark itself isn't OpenAI's.
3Explicit autonomy measurement โ a system that requires constant human correction has limited economic or safety significance regardless of its raw accuracy; benchmarks increasingly try to separate "can do the task" from "can do the task unsupervised."
4Multi-dimensional reporting โ reporting performance depth, task-breadth generality, autonomy level, and reliability as separate axes, rather than compressing everything into one scalar score a single number cannot honestly represent.
โ ๏ธ What's Missing or Uncertain
This article deliberately excludes several claims found during research that could not be independently verified. Reported 2026 scores on Humanity's Last Exam and other benchmarks attributed to model names this site's writers could not confirm exist have been omitted entirely rather than repeated. The claim that OpenAI's original Microsoft partnership included an interim clause tying AGI to a specific profit threshold comes from secondary reporting, not verified primary contract text. Andrew Ng's proposed "Turing-AGI Test," involving giving an agent a real computer and internet access to complete genuine work tasks, is an informal proposal shared via social media, not a published, peer-reviewed framework, and should not be cited as an established benchmark. No entrant has reached a "winning" threshold on any version of the ARC Prize as of the most recent available reporting.
๐ Reference Links
๐ฅ Recommended Videos
๐งญ Closing โ Track the Trend Lines, Not the Finish Line
๐ฏ The Bottom Line
No benchmark in this article claims to measure AGI itself, and the more candid ones โ Humanity's Last Exam included โ say so explicitly. What each of them measures instead is a specific, defensible facet of a concept the field has never agreed how to define: expert knowledge breadth (HLE), skill-acquisition efficiency (ARC-AGI), sustained autonomous task duration (METR), or the ability to do genuine research work (RE-Bench, MLE-bench). The honest way to track "AGI progress" is not to wait for one benchmark to declare victory โ none is designed to, and the ARC Prize has never been won โ but to watch several of these trend lines move together: expert-exam scores climbing, autonomous task duration lengthening, and interactive-environment competence rising from near-zero. That convergence, not any single number, is what the field will eventually look back on as the real signal.