About
A peer-reviewed journal for theoretical AI alignment research. What we publish is set out in the scope and acceptance criteria.
Senior editorial board
Dylan Hadfield-Menell
Associate Professor of EECS at MIT, where he leads the Algorithmic Alignment Group at CSAIL. His research focuses on ensuring AI systems' behavior aligns with the goals of their users and society, including cooperative inverse reinforcement learning, multi-principal assistance games, the off-switch game, multi-agent systems, human-AI teams, and societal oversight of machine learning.
Vanessa Kosoy
Director of AI Research at ALTER and Principal Research Scientist at CORAL, and a former research associate at the Machine Intelligence Research Institute. She leads the learning-theoretic agenda for AI alignment, seeking provable guarantees for safe agents, and originated infra-Bayesianism (with Alexander Appel), a mathematical framework generalizing Bayesian decision theory to handle non-realizability, logical uncertainty, and adversarial environments. Her work spans reinforcement-learning theory, decision theory, and the foundations of embedded agency.
Jan Kulveit
Co-founder and Principal Investigator of the Alignment of Complex Systems (ACS) Research Group, part of the Center for Theoretical Study at Charles University in Prague. Previously a Research Fellow at the Future of Humanity Institute at Oxford. His research focuses on alignment in complex human–AI systems, mathematical theories of hierarchical agency and cooperation, and the psychology of large language models, applying active inference and free-energy methods to model bounded-rational agency. Co-organizes the Human-aligned AI Summer School in Prague.
Seth Lazar
Professor at the Johns Hopkins University School of Government and Policy. He founded the Machine Intelligence and Normative Theory (MINT) Lab, where his research applies moral and political philosophy to AI safety, governance, and institutional design for the AI transition. He has also published widely on the ethics of war and defensive force. His book 'The Algorithmic City: Power, Justice, and AI' is forthcoming from Oxford University Press.
Dan Murfet
Mathematician and Head of Research at Timaeus. Previously a Lecturer (US equivalent: tenured professor) at the School of Mathematics and Statistics at the University of Melbourne. His research focuses on using singular learning theory and developmental interpretability to understand how neural networks learn and generalize. He has also made significant contributions to algebraic geometry and homological algebra.
Tim Rudner
Assistant Professor of Statistical Sciences (Status-Only) at the University of Toronto, a Canada CIFAR AI Chair at the Vector Institute for Artificial Intelligence, and the Chief Scientist at Vijil. He is also a Junior Research Fellow of Trinity College at the University of Cambridge and an Associate Member of the Department of Computer Science at the University of Oxford. His research interests include probabilistic machine learning, AI safety, and AI governance, with a focus on understanding and expanding the statistical foundations of machine learning models, advancing scalable oversight of frontier AI systems, creating trustworthy AI agents, and designing regulatory approaches that enable the effective governance of frontier AI models.
Andrew Saxe
Professor of Theoretical Neuroscience and Machine Learning at the Gatsby Computational Neuroscience Unit and Sainsbury Wellcome Centre at UCL, and a CIFAR Azrieli Global Scholar in the Learning in Machines & Brains programme. His research develops the theory of deep learning and its applications to neuroscience and psychology — including exact analytical solutions for learning dynamics in deep linear networks and a mathematical theory of semantic development.
Benjamin Van Roy
Professor of Electrical Engineering, of Management Science and Engineering, and, by courtesy, of Computer Science at Stanford University. Founder and lead of the Efficient Agent Team at Google DeepMind. His research focuses on reinforcement learning and alignment, with interests including information-theoretic foundations for machine learning, efficient exploration, continual learning, and mathematical models of misalignment risk. He is a Fellow of INFORMS and IEEE and a recipient of the Lanchester Prize.
Advisory board
Scott Aaronson
Professor of Computer Science at UT Austin and founding director of its Quantum Information Center. He researches computational complexity theory and quantum computing, including boson sampling, postselection, and the limits of quantum speedups. As a visiting researcher at OpenAI, he worked on theoretical foundations for AI safety, including AI-output watermarking. He also created the Complexity Zoo and authored Quantum Computing Since Democritus.
Paul Christiano
Founder of the Alignment Research Center (ARC) and, at OpenAI, the originator of reinforcement learning from human feedback (RLHF); he also launched the third-party frontier-model evaluation effort now housed at METR. His research develops scalable oversight and alignment methods — including debate, iterated amplification, and eliciting latent knowledge — for supervising systems whose behavior humans cannot directly check. He has served as Head of AI Safety at the US AI Safety Institute (NIST).
Vince Conitzer
Professor of Computer Science at Carnegie Mellon University, where he directs the Foundations of Cooperative AI Lab (FOCAL). His foundational work in computational social choice, game theory, and mechanism design increasingly addresses AI alignment, focusing on multiagent risks, on how AI systems can represent and aggregate human values, and on broader conceptual issues. He co-authored Moral AI: And How We Get There.
Marcus Hutter
Senior Researcher at DeepMind and Honorary Professor in the Research School of Computer Science at the Australian National University. His research on algorithmic-information-theoretic models of general intelligence unified Solomonoff induction with sequential decision theory in the AIXI framework and related computable approximations. He has also studied reward hacking and value-learning formulations to remove the incentive to manipulate reward signals. Authored Universal Artificial Intelligence and established the €500,000 Prize for Compressing Human Knowledge (the Hutter Prize).
Geoffrey Irving
Co-founder and Chief Scientist of Resolution, a nonprofit applying heavy automation to a portfolio of theoretical and empirical alignment areas. He previously was Chief Scientist at the UK AI Security Institute, led the Scalable Alignment Team at DeepMind and the Reflection Team at OpenAI, and co-led neural network theorem proving work at Google Brain.
Victoria Krakovna
Research Scientist on the AGI Safety & Alignment team at Google DeepMind. She researches deceptive alignment, scheming propensity evaluations, and dangerous capability evaluations. She cofounded the Future of Life Institute.
Jacob Tsimerman
Researcher on the AI safety team at OpenAI and professor of mathematics at the University of Toronto. He was awarded the 2026 Fields Medal for his work in arithmetic geometry, including the proof of the André–Oort conjecture on special points in Shimura varieties. He has also received the Ostrowski Prize and the New Horizons in Mathematics Prize. His recent research concerns risk from AI, including a taxonomy of catastrophic scenarios with Andrew Critch.
Managing editors
Dan MacKinlay
Former Research Scientist at CSIRO’s Data61, Australia’s national information technology laboratory. He is a founding member of LAIR2, the Melbourne AI Safety Hub, and a PIBBSS research resident at the London Initiative for Safe AI. He has written over one million words online about AI, machine learning, philosophy, etc.
Jess Riedel
Physicist and Senior Research Scientist at NTT Research. He primarily researches quantum decoherence and has dabbled in AI alignment as a visiting scholar at the Center for Human-Compatible AI at UC Berkeley. He was recognized by the American Physical Society as an Outstanding Referee, a lifetime award for service as a journal reviewer. He is the lead editor of the Proceedings of ODYSSEY, the 2025 ILIAD conference on AI alignment at Lighthaven.
Follow the journal: Blog X Newsletter