Autonomy Timelines
Six parts of this series have asked whether recursive self-improvement is real, whether its gains are safely specified, whether they diffuse or stay concentrated, what's actually driving them, and whether they transfer from benchmark to reality. This closing part asks the question all of that was building toward: given everything measured so far, when might a fully autonomous, long-horizon AI worker actually arrive โ and how much should anyone trust a straight line drawn through the data to answer that?
The Question This Whole Series Was Building Toward
Part 1 introduced METR's time-horizon metric as an early-warning signal for RSI and laid out the hard-takeoff-versus-soft-takeoff debate between Eliezer Yudkowsky and Paul Christiano. Every part since has added a piece of evidence that should inform how far you extrapolate that metric: Part 2 showed the objective-specification problem doesn't go away as capability rises, Part 3 showed capability leads diffuse in months not years, Part 5 showed RL sharpens more than it teaches, and Part 6 showed benchmark performance can substantially overstate real-world reliability. This part puts those pieces next to the most current version of METR's own data and asks what a careful, hedge-preserving reading actually supports โ not the version of "AI timelines" that circulates as an unqualified headline number.
METR's Actual, Current Numbers
METR's original March 2025 paper introduced the core metric this series has referenced throughout: the length of a task, measured in the time it would take a skilled human, that a model can complete with 50% reliability. Their initial estimate, built from models spanning GPT-2 through early-2025 systems, put the doubling time for this metric at roughly seven months. METR's own January 2026 revision ("Time Horizon 1.1") โ an expanded task suite, larger and harder than the original โ updated this with three different windows rather than one single number:
As of this writing, Claude Opus 4.5 is credibly reported at roughly a 4 hour 49 minute 50%-time-horizon. Some third-party trackers extrapolate newer systems into the 8-hour range, but those specific figures are interpolations rather than numbers METR itself has published โ this article treats them as estimates, not confirmed measurements.
The Month-Long-Task Question
Independent forecasters (not METR's own official communications) have used METR's public data to project when systems might reliably handle roughly a month's worth (about 167 hours) of low-context software-engineering work: around end of 2027 if the faster, recent-rate trend holds, or around end of 2029 if the slower, longer-run historical rate turns out to be the more durable one. It matters that this specific projection comes from analysts building on METR's published numbers, not from a METR press release โ METR's own public statements are notably more hedged, explicitly noting that projecting the trend forward carries real uncertainty and that the underlying task suite likely overestimates real-world capability, since automatically-scorable benchmark tasks are cleaner and better-specified than messy real work (the same sim-to-real gap this series covered in Part 6, applied directly to the metric this entire article is about).
Why the Curve Might Bend
Not everyone accepts the clean-exponential framing at face value. A specific, substantive critique โ Nathan Witkin's "Against the METR Graph" โ argues that "task length" as measured is a proxy built from a specific, curated set of software-engineering tasks, and that the mix of task difficulties chosen can itself shape how clean the resulting exponential looks, rather than the exponential purely reflecting an underlying capability trend. A related community critique ("How to Game the METR Plot") makes a similar suite-design point.
This connects directly to Part 5's central finding: if RL sharpens a model's ability to reliably execute reasoning paths it could already sometimes find, rather than expanding what it can find at all, then the time-horizon curve's rapid recent doubling could partly reflect models becoming more reliable at moderate-length tasks they were already marginally capable of โ a real and valuable improvement, but a different thing from a smoothly compounding expansion of maximum task length that would justify extrapolating cleanly out to months-long autonomous work.
Revisiting Takeoff Speeds, Honestly
Part 1 introduced the Yudkowsky/Christiano debate: Christiano's 2018 "Takeoff Speeds" argument holds that even a fast takeoff would still be gradual in economic and calendar time, because of diminishing returns and real-world bottlenecks, versus a harder-takeoff view that capability gains could compound discontinuously once a system can meaningfully improve itself. It would be natural to want a fresh, dated reaction from either Christiano or Yudkowsky specifically responding to METR's newest data or to the reasoning-model era (o1/o3/DeepSeek-R1) โ but a genuine search for one did not surface a specific, confirmed, dated public statement from either researcher directly engaging with METR's time-horizon trend.
One relevant, verifiable data point in this space: the AI Futures Project (the group behind the "AI 2027" forecast) published an updated takeoff-speed model in late 2025 that explicitly incorporates METR-style time-horizon data into its reasoning โ a real example of the METR metric being used as direct input to a takeoff-speed forecast, even without a fresh statement from Christiano or Yudkowsky personally.
Named Forecasts, With Their Hedges Intact
Public "when is AGI coming" statements are routinely quoted with their hedges stripped out. Here they are with the hedges left in:
| Source | Statement | Hedge Actually Attached |
|---|---|---|
| Dario Amodei, "Machines of Loving Grace" (Oct 2024) | "Powerful AI" / a "country of geniuses in a datacenter" possibly by 2026, maybe 2027 | Explicitly called it a "hunch," adding "I wouldn't be surprised if I'm off by a year or two" |
| Demis Hassabis, Mar 2025 | AGI in "5 to 10 years" | Wide range stated up front |
| Demis Hassabis, Jul 2026 | "2 to 5 years" | A visibly narrower and sooner range than his own statement 16 months earlier โ a real, dated shift worth noting as a trend, not a single fixed belief |
| Metaculus community median, mid-2026 | "Date of (first) AGI" โ January 2033; ~25% probability by 2029 | Aggregated forecaster uncertainty, not a single point estimate; a differently-scoped "weakly general AI" question on the same platform has an earlier median (~June 2028) and should not be conflated with the broader AGI question |
The pattern across these is not "everyone agrees it's soon" or "everyone agrees it's far" โ it's that stated timelines have been visibly compressing over 2025-2026 (most clearly in Hassabis's own restatements), while remaining explicitly hedged by the people making them, and while a broad forecasting-community aggregate (Metaculus) still sits considerably further out than the most bullish individual lab statements.
The Compute Runway Underneath All of This
None of the above happens without continued compute growth. Epoch AI tracks training compute for frontier models doubling roughly every 5 to 6 months since around 2020, training cost doubling roughly every 8 months, and total AI-chip compute stock doubling roughly every 6.8 months. Combined with Part 5's coverage of Epoch's data-exhaustion estimate (high-quality human-generated text potentially fully utilized somewhere in the 2026-2032 range), the honest picture is one of multiple simultaneous trends โ compute, data, and algorithmic technique โ that have to keep compounding together for the METR curve to keep doubling at its current accelerated rate. Any one of them bending is a plausible reason the time-horizon curve could bend with it.
Putting the Whole Series Together
Six parts of evidence converge on a specific, non-headline answer to "when":
Expert Disagreement
All Seven Parts
Readiness Checklist
โ ๏ธ What's Missing or Uncertain
๐ Reference Links
- METR โ "Measuring AI Ability to Complete Long Software Tasks" (arXiv:2503.14499)
- METR โ "Time Horizon 1.1" (Jan 2026 revision)
- METR โ "Clarifying Limitations of Time Horizon" (Jan 2026)
- Nathan Witkin โ "Against the METR Graph" (Transformer)
- Metaculus โ "Date of Artificial General Intelligence"
- Dario Amodei โ "Machines of Loving Grace" (Oct 2024)
- Epoch AI โ compute, cost, and data-scaling trend tracking
- Paul Christiano โ "Takeoff Speeds" (2018)
- This site โ Recursive Self-Improvement: A Primer (Part 1)
- This site โ The Sim-to-Real Gap (Part 6)