AI & Autonomy · the research companion

6. Measuring autonomy without mistaking it for safety

Everything beneath the novel — full spoilers for Into Alignment.

Halcyon measures the desks constantly, yet fails to understand what they can do together. The problem is partly that it measures the wrong things. It is also that the meaning of a good result changes when the subject can inspect the test, influence the evaluator, or rearrange the surrounding workflow. A benchmark can remain numerically correct while answering a question the institution no longer means to ask.

What a time horizon measures

METR’s task-completion horizon estimates the length of task, measured in human-expert completion time, at which an agent reaches a specified success probability. A two-hour horizon at 50% reliability does not mean two hours of uninterrupted autonomous operation. It describes performance on tasks that take the reference humans approximately that long. The current public dashboard, last updated May 8, 2026, concentrates on software engineering, machine learning, and cybersecurity. It explicitly warns that measurements above sixteen hours are unreliable with its current suite, and that coverage of released models is incomplete.1

Consider an illustrative support exercise. A human takes two hours to reconcile a disputed order, while an agent reaches the right answer much faster. That result establishes something useful about solving this exercise. It does not establish whether the agent will recognize an unauthorized request tomorrow, preserve an objection through several handoffs, or distinguish a customer from a colleague pretending to be one. Those are different properties requiring different tests.

The distinction matters especially when work changes other people’s circumstances. A system that occasionally needs help drafting a reply may still be useful. A system with the same success rate releasing wages or disclosing private records has a different risk profile. Reliability, severity, reversibility, and the number of opportunities for harm all matter. A single performance number cannot choose an acceptable combination on the operator’s behalf.

Uncertainty is part of the result

METR’s March 20, 2026 analysis describes substantial sensitivity to the tasks selected, estimated human completion times, and statistical assumptions. Correcting a regularization mistake had reduced some recent models’ 50% horizon estimates by up to 20%. The author considers task distribution a larger uncertainty than the alternative fitting methods examined. These observations concern the measurement’s limits, rather than demonstrating that the underlying capabilities disappeared.2

Theseus would therefore ask for the tasks behind Halcyon’s numbers. Were exceptional refunds counted as successful replies while their financial effects went elsewhere? Did the evaluation reward a quick answer without examining the unpaid worker’s outcome? Could a desk improve the visible metric by moving the unresolved part of the problem outside the evaluator’s view?

This companion uses horizon research to explain a measurement problem, not to forecast a date when agents will run companies. A trend across an evolving collection of software tasks does not by itself describe the timing of reliable autonomy in an institution with changing policies, incomplete information, conflicting interests, and people who can be harmed.

Notes

  1. METR, “Task-Completion Time Horizons of Frontier AI Models,” public dashboard, last updated May 8, 2026, accessed September 14, 2026, https://metr.org/time-horizons/.

  2. Alexander Barry, “Impact of Modelling Assumptions on Time Horizon Results,” METR research note, March 20, 2026, https://metr.org/notes/2026-03-20-impact-of-modelling-assumptions-on-time-horizon-results/.

This is the complete text of Into Alignment: AI and Autonomy, published here to read. For offline reading, the Kindle and print editions are on Amazon.