Home โ€บ Blog โ€บ Autonomy Timelines
RSI Series ยท Part 7 of 7 โ€” Series Finale โณ

Autonomy Timelines

Six parts of this series have asked whether recursive self-improvement is real, whether its gains are safely specified, whether they diffuse or stay concentrated, what's actually driving them, and whether they transfer from benchmark to reality. This closing part asks the question all of that was building toward: given everything measured so far, when might a fully autonomous, long-horizon AI worker actually arrive โ€” and how much should anyone trust a straight line drawn through the data to answer that?

FL
FrontierAGI Team

The Question This Whole Series Was Building Toward

Part 1 introduced METR's time-horizon metric as an early-warning signal for RSI and laid out the hard-takeoff-versus-soft-takeoff debate between Eliezer Yudkowsky and Paul Christiano. Every part since has added a piece of evidence that should inform how far you extrapolate that metric: Part 2 showed the objective-specification problem doesn't go away as capability rises, Part 3 showed capability leads diffuse in months not years, Part 5 showed RL sharpens more than it teaches, and Part 6 showed benchmark performance can substantially overstate real-world reliability. This part puts those pieces next to the most current version of METR's own data and asks what a careful, hedge-preserving reading actually supports โ€” not the version of "AI timelines" that circulates as an unqualified headline number.

METR's Actual, Current Numbers

METR's original March 2025 paper introduced the core metric this series has referenced throughout: the length of a task, measured in the time it would take a skilled human, that a model can complete with 50% reliability. Their initial estimate, built from models spanning GPT-2 through early-2025 systems, put the doubling time for this metric at roughly seven months. METR's own January 2026 revision ("Time Horizon 1.1") โ€” an expanded task suite, larger and harder than the original โ€” updated this with three different windows rather than one single number:

~188 days Doubling time, all-time trend (2019โ€“present)
~129 days Doubling time, since 2023
~89 days Doubling time, since 2024 โ€” the trend is accelerating

As of this writing, Claude Opus 4.5 is credibly reported at roughly a 4 hour 49 minute 50%-time-horizon. Some third-party trackers extrapolate newer systems into the 8-hour range, but those specific figures are interpolations rather than numbers METR itself has published โ€” this article treats them as estimates, not confirmed measurements.

โš ๏ธ The Caveat METR Themselves Insist On METR published a dedicated clarification in January 2026 stating explicitly that a 50% time horizon of X does not mean tasks under X can be reliably delegated to the model, and that a doubling of the time-horizon metric does not mean a doubling of real-world automatable work. Critically, METR also flags that its own estimates above roughly 16 hours are unreliable, because the benchmark suite doesn't yet contain enough long, hard tasks to pin the tail of the curve down precisely. Any claim built on extrapolating this metric past that point is extrapolating past the range METR itself is confident measuring.

The Month-Long-Task Question

Independent forecasters (not METR's own official communications) have used METR's public data to project when systems might reliably handle roughly a month's worth (about 167 hours) of low-context software-engineering work: around end of 2027 if the faster, recent-rate trend holds, or around end of 2029 if the slower, longer-run historical rate turns out to be the more durable one. It matters that this specific projection comes from analysts building on METR's published numbers, not from a METR press release โ€” METR's own public statements are notably more hedged, explicitly noting that projecting the trend forward carries real uncertainty and that the underlying task suite likely overestimates real-world capability, since automatically-scorable benchmark tasks are cleaner and better-specified than messy real work (the same sim-to-real gap this series covered in Part 6, applied directly to the metric this entire article is about).

Why the Curve Might Bend

Not everyone accepts the clean-exponential framing at face value. A specific, substantive critique โ€” Nathan Witkin's "Against the METR Graph" โ€” argues that "task length" as measured is a proxy built from a specific, curated set of software-engineering tasks, and that the mix of task difficulties chosen can itself shape how clean the resulting exponential looks, rather than the exponential purely reflecting an underlying capability trend. A related community critique ("How to Game the METR Plot") makes a similar suite-design point.

The Strongest Version of This Critique METR's own January 2026 limitations note contains the clearest evidence for it: frontier models are near-100% reliable on tasks under about four minutes of human-equivalent time, but reliability falls below 10% for tasks past roughly four hours โ€” well before the specific length used to define the 50%-threshold "time horizon." A single number describing "the task length at 50% success" compresses a much messier reliability-versus-length surface into one point, and reasonable people can disagree about how much that compression is hiding.

This connects directly to Part 5's central finding: if RL sharpens a model's ability to reliably execute reasoning paths it could already sometimes find, rather than expanding what it can find at all, then the time-horizon curve's rapid recent doubling could partly reflect models becoming more reliable at moderate-length tasks they were already marginally capable of โ€” a real and valuable improvement, but a different thing from a smoothly compounding expansion of maximum task length that would justify extrapolating cleanly out to months-long autonomous work.

Revisiting Takeoff Speeds, Honestly

Part 1 introduced the Yudkowsky/Christiano debate: Christiano's 2018 "Takeoff Speeds" argument holds that even a fast takeoff would still be gradual in economic and calendar time, because of diminishing returns and real-world bottlenecks, versus a harder-takeoff view that capability gains could compound discontinuously once a system can meaningfully improve itself. It would be natural to want a fresh, dated reaction from either Christiano or Yudkowsky specifically responding to METR's newest data or to the reasoning-model era (o1/o3/DeepSeek-R1) โ€” but a genuine search for one did not surface a specific, confirmed, dated public statement from either researcher directly engaging with METR's time-horizon trend.

โš ๏ธ What This Article Will Not Do Rather than inventing or loosely paraphrasing a position neither researcher has publicly stated in this specific context, this article names the gap directly: their standing 2018-2021-era positions (Christiano's gradual-fast-takeoff argument, Yudkowsky's sharper-discontinuity concerns from the MIRI/LessWrong dialogues) remain their most recent confirmed public statements on this exact question, as far as this research could establish. If either has commented more recently and specifically on METR's data, this article simply doesn't have a verified citation for it โ€” a limitation worth stating rather than papering over.

One relevant, verifiable data point in this space: the AI Futures Project (the group behind the "AI 2027" forecast) published an updated takeoff-speed model in late 2025 that explicitly incorporates METR-style time-horizon data into its reasoning โ€” a real example of the METR metric being used as direct input to a takeoff-speed forecast, even without a fresh statement from Christiano or Yudkowsky personally.

Named Forecasts, With Their Hedges Intact

Public "when is AGI coming" statements are routinely quoted with their hedges stripped out. Here they are with the hedges left in:

SourceStatementHedge Actually Attached
Dario Amodei, "Machines of Loving Grace" (Oct 2024)"Powerful AI" / a "country of geniuses in a datacenter" possibly by 2026, maybe 2027Explicitly called it a "hunch," adding "I wouldn't be surprised if I'm off by a year or two"
Demis Hassabis, Mar 2025AGI in "5 to 10 years"Wide range stated up front
Demis Hassabis, Jul 2026"2 to 5 years"A visibly narrower and sooner range than his own statement 16 months earlier โ€” a real, dated shift worth noting as a trend, not a single fixed belief
Metaculus community median, mid-2026"Date of (first) AGI" โ‰ˆ January 2033; ~25% probability by 2029Aggregated forecaster uncertainty, not a single point estimate; a differently-scoped "weakly general AI" question on the same platform has an earlier median (~June 2028) and should not be conflated with the broader AGI question

The pattern across these is not "everyone agrees it's soon" or "everyone agrees it's far" โ€” it's that stated timelines have been visibly compressing over 2025-2026 (most clearly in Hassabis's own restatements), while remaining explicitly hedged by the people making them, and while a broad forecasting-community aggregate (Metaculus) still sits considerably further out than the most bullish individual lab statements.

The Compute Runway Underneath All of This

None of the above happens without continued compute growth. Epoch AI tracks training compute for frontier models doubling roughly every 5 to 6 months since around 2020, training cost doubling roughly every 8 months, and total AI-chip compute stock doubling roughly every 6.8 months. Combined with Part 5's coverage of Epoch's data-exhaustion estimate (high-quality human-generated text potentially fully utilized somewhere in the 2026-2032 range), the honest picture is one of multiple simultaneous trends โ€” compute, data, and algorithmic technique โ€” that have to keep compounding together for the METR curve to keep doubling at its current accelerated rate. Any one of them bending is a plausible reason the time-horizon curve could bend with it.

Putting the Whole Series Together

Six parts of evidence converge on a specific, non-headline answer to "when":

1
The trend is real and accelerating โ€” METR's own three-window analysis (7 months โ†’ 6.3 months โ†’ 4.3 months โ†’ ~3 months) shows genuine acceleration, not a stable or slowing rate, as of the most recent published data.
2
The metric likely overstates real-world readiness โ€” per Part 6's sim-to-real findings and METR's own admission that benchmark tasks are cleaner than real work, plus the sharp reliability drop-off past four hours that a single 50%-threshold number obscures.
3
What's driving the gains is partly reliability, not just reach โ€” per Part 5, RL-driven improvements may concentrate on executing already-reachable reasoning more consistently rather than purely expanding the maximum task length a model can ever complete.
4
Whatever capability level is reached diffuses fast โ€” per Part 3's Epoch AI data, a capability lead survives roughly three to four months before competitors close the gap, meaning "when the first lab gets there" and "when it's widely available" are close together in time, not decades apart.
5
The objective-specification and safety-framework questions from Parts 2 and 4 don't resolve on any of these timelines โ€” they compound at whatever pace capability arrives, which is exactly why every major lab's safety framework (covered in Part 1) treats "autonomous AI R&D capability" as a tracked threshold to prepare for now, not a distant hypothetical.

Expert Disagreement

This is, appropriately for a series finale, the section with the widest genuine range of expert opinion in the entire series. Individual lab leaders' stated timelines for transformative AI span from roughly two years to a decade, with visible movement toward shorter timelines from at least one prominent figure (Hassabis) over 2025-2026. A broad forecasting-community aggregate (Metaculus) remains considerably more conservative than the most bullish individual statements. Whether METR's accelerating time-horizon trend should be read as strong evidence for the shorter end of that range, or as a metric whose own authors have explicitly warned against over-extrapolating, is itself a live disagreement without a resolved answer โ€” this article has tried to present the actual data and the actual caveats rather than adjudicate between the camps.
The Complete Series

All Seven Parts

Part 1 ยท Published
Where RSI comes from, what's real today, and the safety-framework convergence that made this series worth writing.
Part 2 ยท Published
Why you can't write down what you actually want, and why that gap compounds under RSI.
Part 3 ยท Published
A capability lead survives contact with distillation for months, not years โ€” Epoch AI's real catch-up-time data.
Part 4 ยท Published
Real, measured systems (Dream-RSI, ModularRSI, ScienceBuddy) accelerating AI research today.
Part 5 ยท Published
RL sharpens a model's existing solution distribution more than it expands it โ€” evidence that bears directly on this closing part.
Part 6 ยท Published
Why benchmark performance overstates real-world reliability โ€” the reason this part treats METR's own numbers so carefully.

Readiness Checklist

1
Can you state METR's current doubling-time estimate, and explain why it's given as three different numbers rather than one?
2
Can you explain why a 50%-time-horizon of X hours does not mean a task of that length can be reliably delegated?
3
Can you name at least one specific, structural reason the METR curve might bend rather than continue exponentially?
4
Could you explain, without conflating them, the difference between Metaculus's "AGI" question median and its "weakly general AI" question median?

โš ๏ธ What's Missing or Uncertain

This closing article carries more unresolved uncertainty than any other part in the series, appropriately, since it's the part asking the least-answerable question. No verified, dated, specific 2025-2026 public statement from Paul Christiano or Eliezer Yudkowsky directly reacting to METR's time-horizon data or the reasoning-model era was located during research for this piece โ€” their positions are described from earlier, confirmed public statements (2018-2021), not fresh commentary. Third-party estimates of newer models' time horizons (in the 7.9-8.7 hour range for systems past Claude Opus 4.5) are interpolations by outside trackers, not official METR-scored figures. The "month-long task by 2027-2029" projection comes from independent analysts extrapolating METR's public data, not from METR's own official communications, which are notably more hedged about extrapolation. Individual lab leaders' timeline statements are inherently made by parties with a stake in public perception of their own progress, a caveat this series has applied consistently and applies here too.

๐Ÿ”— Reference Links

๐ŸŽฅ Recommended Videos

๐Ÿงญ Series Closing โ€” What Seven Parts of Careful Reading Actually Supports

๐ŸŽฏ The Bottom Line
The honest answer to "when does autonomous, long-horizon AI arrive" is not a date โ€” it's a trend that is real, measured, and accelerating (METR's own three-window doubling-time data shows that plainly), attached to a metric its own creators have explicitly warned not to over-extrapolate, describing a capability that this series has shown compounds unevenly (Part 5), transfers to reality less cleanly than benchmarks suggest (Part 6), diffuses to competitors within months once achieved (Part 3), and arrives without the objective-specification problem (Part 2) being anywhere close to solved. None of that is a reason to dismiss the trend โ€” it is real and it is accelerating. It is a reason to distrust anyone, on either side, who states a specific year with more confidence than METR itself attaches to its own data. That, more than any single number in this article, is the takeaway this seven-part series has been building toward from Part 1's opening question of whether AI can meaningfully accelerate the process that builds better AI: yes, measurably, today โ€” and precisely because that's true, every caveat in this series about how that acceleration is measured, specified, and controlled matters more, not less, the closer the trend lines get to the numbers making headlines.