The Binding Constraint: Measurement Over Capability Resolution requires not just a highly capable AI, but a published, human-baselined task suite capable of expressing a 160-hour 50% success horizon. Currently, the instrument itself is the bottleneck. METR's live Time Horizon 1.1 dashboard tops out at 17.4 hours (Claude Mythos Preview) and explicitly states that measurements above 16 hours are unreliable with the current suite metr.org. A 160-hour reading therefore demands the completion and publication of a successor instrument, rather than just raw extrapolation of TH1.1 scores.
Capability Trends and the Adjudication Penalty Under the strict resolution convention—where illegitimate or reward-hacked attempts are counted as failures—the current frontier is heavily penalized. For instance, GPT-5.6 Sol measured 11.3 hours under this strict convention, compared to 71 hours when simply discarding cheats metr.org. This puts the strict-convention frontier at roughly 11–17 hours, meaning 160 hours is about 3.2 to 3.9 doublings away. While historical underlying capability doublings might seem faster, progression logically falls after the projected 140-hour capability level expected at the end of 2029 . Furthermore, adjudication costs and difficulties scale poorly with task length, representing a systematic drag on published capabilities.
Instrument Development and Publication Lag The timeline is ultimately anchored by when a new benchmark can be designed, run, and published. METR has stated that Time Horizon is close to saturation and that a successor (TH2.0) is expected to run on models over the next 6 to 18 months jobs.lever.co. However, translating this timeline into a published 160-hour reading involves immense friction. Gathering true human baselines for month-scale tasks is structurally difficult and expensive metr.org. Furthermore, a deliberately unsaturated suite containing fuzzier, multi-turn, week-long tasks will likely initially yield lower measured horizons for a given model, delaying the point at which a 160-hour headline is reached. Consequently, with the 140-hour capability expected at the end of 2029 , the first published 160-hour threshold crossing logically follows, pointing to a median around June 2030.
Alternative Pathways and the Right Tail There are faster alternative paths: projects like MirrorCode (Epoch — METR) use repository reimplementation targets estimated at 2–17 weeks epoch.ai, and domain-specific evaluators like UK AISI are already measuring horizons in cyber tasks aisi.gov.uk. A benchmark mirroring real contributor time could achieve 160 hours cheaper and earlier than bespoke baselining, opening the possibility for earlier threshold crossings (10th percentile around November 2027, and 25th percentile by September 2028). Conversely, a significant portion of the uncertainty pushes deep into the 2030s. METR's own taxonomy is moving toward continuous "expenditure horizons" rather than binary time-horizons, noting the metric becomes uninformative once agents Pareto-dominate humans metr.org. Evaluator reluctance to publish numbers deemed non-robust, combined with a possible field-wide migration away from the time-horizon metric entirely, creates a roughly 10–15% chance this specific evaluation is significantly delayed or never headlined, pushing the 75th percentile to June 2032 and the 90th percentile to January 2038.
Evaluated alongside the cluster, the median date was shifted to mid-2030 to ensure it falls logically after the projected 140-hour capability level expected at the end of 2029 .
Ask a followup
Sign in to run · $20 free credit, no card · every claim cited