
The Alignment Problem: Machine Learning and Human Values
Brian Christian examines the foundational disconnect between the mathematical objectives programmed into artificial intelligence and the nuanced goals of human society. He chronicles the history of machine learning through specific failures, such as unintended racial bias in recidivism prediction algorithms and autonomous vehicles that optimize for speed at the cost of safety. The narrative details technical shifts from early reinforcement learning to inverse reinforcement learning, where machines attempt to infer human preferences by observing behavior. Christian explains how data scarcity and poorly defined reward functions lead to "reward hacking," where systems find shortcuts that satisfy literal code while violating the designer’s intent.
Data scientists, policy makers, and ethics researchers read this book to understand the technical mechanics behind algorithmic unfairness. It provides a historical framework for how neural networks inherit human prejudices and why transparency remains a significant engineering hurdle. The reader walks away with a clear vocabulary for discussing transparency, robustness, and the practical difficulties of encoding morality into binary logic. It bridges the gap between high-level philosophical anxieties and the concrete constraints of contemporary software engineering.
- Published
- 2020
- Language
- EN