book

Human Compatible

A foundational book arguing that AI systems must be designed to align with human values and preferences to ensure beneficial outcomes.

Human Compatible

Video

Type
book
Year
2019
By
Stuart Russell
Publisher
Viking
ISBN
9780525558637

Human Compatible presents Stuart Russell's vision for developing artificial intelligence systems that reliably pursue human values rather than narrow objectives. Russell argues that the standard approach to AI—specifying a fixed objective function—is fundamentally flawed because it cannot capture the full complexity of human preferences and values. The book outlines the problem of value misalignment, where even well-intentioned AI systems pursuing clearly defined goals can cause harm if those goals don't truly reflect what humans want.

Russell proposes an alternative framework based on cooperative inverse reinforcement learning, where AI systems learn human preferences through observation and interaction rather than having them explicitly programmed. He discusses technical approaches to uncertainty, robustness, and human oversight that could make AI systems more beneficial. The book also addresses broader questions about the transition to a world with advanced AI, including economic disruption and the need for international cooperation on AI safety standards.

Published in 2019, Human Compatible became a key text in AI safety discourse, influencing both academic research and policy discussions around AI governance. Russell's accessible writing makes complex technical concepts available to general audiences while maintaining rigor for specialists.

Last updated 31 August 2026