Why agency?
We define agency to mean autonomous, goal-directed behaviour, and related phenomena. We believe understanding such phenomena is essential for addressing the AI alignment problem — the problem of ensuring that powerful, autonomous AI pursues goals that are compatible with human (and animal) flourishing.
With AI capabilities rapidly progressing, now is a crucial moment to develop a scientifically-grounded theory of which types of agents are robustly safe. In particular, we seek a thermodynamics of agency, a theory of macroscopic principles which do not appeal to specific details of the architectures of neural networks or brains.