METR's time-horizon milestones, GPT-2 to Claude Opus 4.6
METR, an AI evaluation nonprofit, measures the length of task — in the minutes a skilled human professional would need — that a model can complete on its own with a 50% success rate. In February 2019, GPT-2 managed about 2 seconds of task length, roughly what it takes to read a single sentence. By February 2026, Claude Opus 4.6 reached about 718 minutes, or roughly 12 hours — a full workday a person could hand off and walk away from with even odds it gets finished unsupervised.
Between those two points sit six more milestones, each mapped to a task a person would recognize rather than a bare number: GPT-4 answering a quick question or fixing a one-line bug at 4 minutes in 2023, Claude 3.5 Sonnet handling a short coding task at 35 minutes in 2024, and o3 covering most of a morning's focused work at 110 minutes by April 2025.
The dataset behind this piece is eight confirmed anchor points spanning seven years, not a complete census of every model released in that window — several 2023-2026 releases could not be sourced with confidence and are left out rather than guessed at.
! The Milestones That Define the Climb:
•
From a Sentence to a Bug Fix, 2019–2023:GPT-2 (Feb 2019) managed roughly 2-second tasks — reading a single sentence. GPT-4 (Mar 2023) reached about 4 minutes — answering a quick question or fixing a one-line bug.
•
From a Coding Task to Half a Workday, 2024–2025:Claude 3.5 Sonnet hit ~35 minutes of real coding work in late 2024. By August 2025, GPT-5 reached 2h17m — a solid half-day chunk of work.
•
A Full Workday, Unsupervised, by 2026:Claude Opus 4.6 reached roughly 718 minutes (~12 hours) in February 2026 — a task long enough to occupy a skilled professional for a full working day.
? The Numbers Behind the Story:
•
GPT-2, Feb 2019:2 seconds — read a sentence
•
GPT-4, Mar 2023:~4 minutes — fix a one-line bug
•
o3, Apr 2025:~110 minutes — most of a morning
•
Claude Opus 4.6, Feb 2026:~12 hours (CI 5.3-65.8h) — a full workday
The headline is autonomy, not intelligence: how long a task a frontier model can be trusted with, unsupervised, before it needs a human to step back in. Told through plain milestones — a sentence, a bug fix, a coding task, a morning, a workday — the climb from 2019 to 2026 is easy to picture even without a single doubling-time calculation. The dataset behind it is honest about its limits, too: eight confirmed points, wide uncertainty bands at the long end, and several 2023-2026 releases left out because a confident minute-level figure could not be sourced for them.
Continue reading below for the full detailed article →
Overview
Why 'How Smart' Was Never a Real Question
"How smart is this AI model?" is not a question a dataset can answer, because "smart" is not a unit. METR, an AI evaluation nonprofit, built something narrower and more useful instead: the 50% time horizon, defined as the length of task — measured in the minutes a skilled human professional would need — that a given AI agent can complete on its own about half the time. It is not a proxy for general intelligence. It is a direct measure of autonomy: how long a model can be trusted to work without a human stepping in. This piece follows eight confirmed data points from GPT-2 in February 2019 to Claude Opus 4.6 in February 2026, translating each one into a task a person would recognize — a sentence, a bug fix, a full workday — rather than a bare number.
Three Milestones in How Long AI Can Work Alone
One number marks where this story starts, one marks the point AI crossed from minutes into hours, and one marks where it stands today — a full workday, unsupervised, in February 2026.
2 Seconds: Where the Climb Starts
2 sec
GPT-2, released in February 2019, could complete roughly 2-second tasks with 50% reliability — about as long as it takes a skilled person to read a single sentence. It is the floor of this entire dataset and the baseline every later model is measured against, the starting point for a climb that would eventually reach a full unsupervised workday seven years later.
110 Minutes: A Morning's Work, April 2025
110 min (~1h 50m)
o3, released by OpenAI in April 2025, reached a 110-minute time horizon — enough for most of a morning's focused work, up from just 4 minutes two years earlier with GPT-4. It's the midpoint of this dataset's climb, the model that first pushed the time horizon convincingly past the one-hour mark and into multi-hour territory the rest of 2025 would build on.
12 Hours: A Full Workday, Unsupervised
718 min (~12h)
Claude Opus 4.6, added to METR's leaderboard on February 20, 2026, reached roughly 718 minutes — about 12 hours, or a full workday a skilled professional could hand off and walk away from with even odds it finishes unsupervised. METR's own confidence interval for that estimate spans 5.3 to 65.8 hours, a reminder that the newest, longest-horizon points carry real, wide uncertainty even as the headline number keeps climbing.
Claude Opus 4.6 by the Numbers
The Point Estimate: ~12 Hours 718 minClaude Opus 4.6's 50% time horizon lands at roughly 718 minutes, or about 12 hours — the length of task, in skilled-human time, it can complete unsupervised with even odds of success.
The Jump From GPT-2: ~21,800x 21,800 x since 2019Measured against GPT-2's 2-second time horizon in February 2019, Claude Opus 4.6's roughly 718-minute (43,080-second) score is a jump of about 21,800 times over seven years.
The Uncertainty Band: 5.3–65.8 Hours 5.3-65.8 hours (95% CI)METR's own 95% confidence interval for Claude Opus 4.6 spans 5.3 to 65.8 hours — a wide range that reflects how few very long tasks exist to test against at that end of the scale.
The Metric, Explained
A Difficulty Ceiling, Not a Stopwatch
This is not a measure of how long an AI takes to think. It's a difficulty scale, measured in human time. METR gives each model a large set of test tasks — some short enough that a skilled person finishes in seconds, some long enough to take a full day — and finds the task length where the model starts failing about half the time. That length, expressed in the minutes a skilled human professional would need, is the model's time horizon.
So when Claude Opus 4.6 is described as having a roughly 12-hour time horizon, the claim is specific: give it something that would take a skilled person about 12 hours, and it succeeds roughly half the time. Give it something a person would need two full days for, and it will usually fail. That's a ceiling on the task difficulty a model can be trusted with unsupervised — not a measurement of how fast the model itself works.
How It's Measured
How METR Turns Task Difficulty Into Minutes
METR builds the time horizon from three task suites: HCAST (real software engineering work), RE-Bench (machine learning research tasks), and SWAA (short, atomic actions used to anchor the low end of the scale). Every task in those suites also has a human baseline — how long a paid contractor with the relevant skill actually took to finish it. A model's 50% time horizon is the task length, on that human-minutes scale, where the model's success rate crosses 50%.
The design choice that matters most is the human baseline. It converts an abstract pass/fail benchmark score into something with a real unit: minutes of skilled human work. That is what makes GPT-2's 2 seconds and Claude Opus 4.6's roughly 12 hours comparable numbers on the same axis, rather than two different scores on two different scales. METR published the full methodology in "Measuring AI Ability to Complete Long Software Tasks" (Kwa et al., arXiv:2503.14499, March 2025), later presented at NeurIPS 2025, and has since expanded the leaderboard with a January 2026 update called Time Horizon 1.1.
The Early Climb
From Reading a Sentence to Fixing a Bug
GPT-2, released in February 2019, managed roughly 2 seconds of task length at 50% reliability — about as long as it takes to read a single sentence, and essentially the floor of this entire dataset. Four years later, GPT-4 (the March 2023 checkpoint) reached about 4 minutes — enough to answer a quick factual question or fix a one-line bug — a jump of two orders of magnitude that reflects the leap from a text predictor to a genuinely task-capable model.
Claude 3.5 Sonnet (New) followed in October 2024 at about 35 minutes — a short, well-defined coding task a developer could hand off and expect back, more or less finished, without checking in halfway through. It's the first model in this dataset that reads less like a demo of what AI can do and more like a tool someone could actually delegate real, bounded work to.
2025: A Crowded Year
Four Models, Four Steps Toward a Full Day
February 2025 landed two milestones within days of each other: GPT-4.5's preview checkpoint reached about 30 minutes — a short, well-defined task like writing a small script or summarizing a document — and Claude 3.7 Sonnet reached 50 minutes, roughly a focused, coffee-break-length task. Neither model was dramatically ahead of the other; together they marked the point where 'AI task length' stopped being measured in single-digit minutes.
o3, released by OpenAI in April 2025, reached about 110 minutes — close to two hours, and enough to work through most of a morning's focused task before a human would need to check back in.
GPT-5 followed in August 2025 at a 2h17m (137-minute) point estimate — a solid half-day chunk of work. METR's own 95% confidence interval for that figure runs from 65 minutes to 4 hours 25 minutes, a wide band that is typical of estimates this far out on the scale, where fewer test tasks exist to pin the number down precisely.
The Newest Point
Claude Opus 4.6 and the Jump to a Full Workday
Claude Opus 4.6, added to METR's leaderboard on February 20, 2026, is the highest point in this dataset by a wide margin: a 50% time horizon of roughly 718 minutes, or about 12 hours. That is more than five times o3's score from ten months earlier, and it means a model can now be trusted, at 50% reliability, to work through a task that would occupy a skilled professional for the better part of a working day.
It is also the least certain point on the chart. METR's own published confidence interval for Opus 4.6 runs from about 5.3 hours to 65.8 hours — a band wide enough that the true value could plausibly sit anywhere from "a long afternoon" to "most of a work week." That is not a flaw unique to this model; it is a structural feature of measuring long time horizons, where very few tasks exist that take that long to complete, so the estimate rests on a thin slice of the test suite.
A separate compilation of the same underlying data put Opus 4.6 closer to 14.5 hours rather than 12. Both figures fall comfortably inside the published 5.3-65.8 hour interval, so this is less a factual contradiction than a reminder of how much room that interval leaves for disagreement about the exact point estimate.
Limits
What Eight Anchor Points Cannot Tell You
This dataset is eight confirmed points across seven years, not a complete model census. METR's own leaderboard has added several models this piece deliberately leaves out — GPT-4o, o1, Claude Opus 4.5, Claude Sonnet 4.5, Gemini 3 Pro, Grok 4, and the GPT-5.1 through 5.4 line among them — because a minute-level point estimate for each could not be sourced from a reference solid enough to stand behind in this pass. Filling those gaps would meaningfully thicken the middle and right of the timeline, and is flagged as follow-up work rather than guessed at here.
GPT-3 and GPT-3.5 are absent for a related reason: METR itself dropped them from the current leaderboard, citing the need for significant changes to the tool-calling scaffold to evaluate them fairly, and no primary-sourced minute-level figure for either could be confirmed independently.
One further wrinkle worth naming: METR's own sources disagree slightly on Claude 3.5 Sonnet (New), with a preliminary evaluation report putting it at about 35 minutes and the later time-horizon paper placing the broader late-2024 model cohort closer to 40 minutes. Both figures are METR-sourced and the gap likely reflects different task-suite versions rather than an error; this piece uses the 35-minute figure throughout, as the more specific, model-level estimate.
The Full Dataset
All 8 Confirmed Time-Horizon Milestones, 2019-2026
Every model in this dataset, with its 50% time horizon, its plain-language equivalent, and the published confidence interval where one exists. Point estimates for GPT-5 and Claude Opus 4.6 carry the widest uncertainty bands, reflecting how few very long tasks exist to test against at that end of the scale.
GPT-2
OpenAI
Feb 2019
2 seconds
Read a single sentence
—
Primary
GPT-4 (0314)
OpenAI
Mar 2023
~4 minutes
Answer a quick question, or fix a one-line bug
—
Primary
Claude 3.5 Sonnet (New)
Anthropic
Oct 2024
~35 minutes
A short, well-defined coding task
—
Primary
Claude 3.7 Sonnet
Anthropic
Feb 2025
50 minutes
A focused, coffee-break-length task
—
Primary
GPT-4.5 (preview)
OpenAI
Feb 2025
~30 minutes
A short, well-defined task — a small script, a document summary
—
Primary
o3
OpenAI
Apr 2025
110 minutes (~1h50m)
Most of a morning's focused work
—
Primary
GPT-5
OpenAI
Aug 2025
2h 17m
A solid half-day chunk of work
65 min-4h25m
Secondary (corroborated)
Claude Opus 4.6
Anthropic
Feb 2026
~12 hours
A full workday, unsupervised
5.3h-65.8h
Secondary (derived)
Confidence tiers reflect sourcing, not model quality: "Primary" is a figure drawn directly from METR's own published paper, evaluation report, or official post; "Secondary" is a figure derived from or corroborating METR's published data via independent analysis. Nine additional 2023-2026 model releases could not be sourced to a confident minute-level figure and are omitted rather than estimated.
Verdict
The Real Story Is Autonomy, Not Intelligence
Read as a set of milestones rather than a statistic, the climb is easy to picture: a sentence in 2019, a one-line bug fix in 2023, a real coding task in 2024, a morning's work in April 2025, a half-day chunk by August 2025, and a full unsupervised workday by February 2026. Each step is a task a person would recognize, which is exactly the point of measuring 'time horizon' in human minutes rather than an abstract benchmark score.
It's worth being just as clear about what this dataset doesn't settle. Eight confirmed anchor points is a strong skeleton, not a full census — several 2023-2026 releases are missing because a confident minute-level figure couldn't be sourced for them, and the two newest points, GPT-5 and Claude Opus 4.6, carry the widest published uncertainty bands in the whole series.
What comes next is a question of scale, not a new kind of question: will the frontier keep moving from 'a full workday' toward 'a full week,' and how will METR's own methodology hold up as it tries to measure tasks long enough that almost no test suite has enough of them to test against.
The data behind this story comes from METR (Model Evaluation & Threat Research), a nonprofit AI evaluation organization, via its March 2025 paper "Measuring AI Ability to Complete Long Software Tasks" (Kwa et al., arXiv:2503.14499, later presented at NeurIPS 2025), its January 2026 Time Horizon 1.1 update, its official model evaluation reports, and its public time-horizon leaderboard. Individual model figures for GPT-4.5 and Claude Opus 4.6 additionally draw on METR's own social posts and independent analyses of METR's published data where a direct paper citation was not available. The dataset is publicly published by METR for research and journalistic use, and full credit for designing the methodology, running the evaluations, and maintaining the leaderboard goes to METR's research team.
FactsFigs reviews, cleans, and cross-checks every source dataset before shaping it into a data story. Each visualization is created and designed in FactsFigs Design Studio — an internal tool developed and owned by FactsFigs — and is the original work of a FactsFigs author, not an AI-generated copy of any existing graphic. Individual assets within a visual may or may not be produced with AI tools, but the design of the visual itself is solely FactsFigs' own.
Figures are estimates at the time of publication, provided for information only — nothing here is financial advice or a guarantee of accuracy.