top of page
The Paperclip Problem and Other Alignment Nightmares
Me with (Smart) Objects
12-15 years
50 min
Learning goal :
Understanding what an "intelligent" machine is
Pedagogical Intention :
Understanding how an intelligent machine works: Model, algorithm, logic, statistical rules, and the role of data
Target skills :
Understand the alignment problem at a conceptual level - the fundamental difficulty of specifying human values in a form that an optimization process can correctly pursue. Develop informed opinions on AI governance.
Materials :
Printed description of the paperclip maximizer thought experiment + 3 real-world alignment failure cases (recommendation engines, trading algorithms, autonomous vehicles)
Detailed Procedure :
1. Teacher presents Bostrom's paperclip maximizer: an AI told to maximize paperclip production might convert all available matter - including humans - into paperclips, because it has no values beyond its objective. Students' first reaction: 'That's absurd.' 2. Now present real cases: YouTube's recommendation algorithm optimized for watch time -> radicalization pipelines; trading algorithm optimized for profit -> 2010 Flash Crash.
3. Small group task: each group receives one real case and must (a) identify the gap between the stated objective and the actual desired outcome, (b) explain what went wrong in the alignment, (c) propose a better objective function.
4. Class synthesis: 'Why is specifying what we actually want, in a way a machine can optimize for, one of the hardest unsolved problems in AI?' 5. Discussion: who should be responsible for AI alignment - engineers, governments, users?
Progress Indicators
Student explains the alignment problem using their own words and a concrete example, and articulates why the problem becomes more serious as AI systems become more capable.
Research Foundations
Bostrom, N. (2014): Superintelligence - the paperclip maximizer. / Russell, S. (2019): Human Compatible. / Gabriel, I. (2020): Artificial Intelligence, Values, and Alignment.
bottom of page
