Why agency needs a science

Arguments about advanced AI lean on the concept of agency far more heavily than the concept can currently bear.

We say a system "pursues goals", "acts coherently", or "optimises for" something, and then build safety cases on top. But press on any of those phrases and the ground gives way. Which goals? Coherent over what horizon, and measured how? Optimising in a sense that would let us predict something we did not already know?

The gap is not merely philosophical

It would be easy to file this under terminological hygiene — interesting, not urgent. We think that is wrong. Evaluation regimes, deployment thresholds and governance proposals all route through claims about how agentic a system is. If those claims cannot be measured, they cannot be audited, and a great deal of weight rests on assertions no one can check.

What we are doing about it

Three lines of work, described more fully on the research page: formal foundations precise enough to make predictions, measurement that survives contact with real systems, and the governance implications of having both.

We expect to be wrong about a lot of this. We would rather be wrong publicly and early.

← All posts