DeepMind's David Silver explains how AlphaGo, AlphaZero, and MuZero used self-play reinforcement learning to master games and discover superhuman creativity.

David Silver: Leader of DeepMind's reinforcement learning research group and the lead researcher on AlphaGo and AlphaZero, who also co-led AlphaStar and MuZero. He is one of the central figures behind modern deep reinforcement learning.
David Silver traces his path from programming a BBC Micro at age seven and building games to a PhD applying reinforcement learning to the game of Go. He explains the core of reinforcement learning, why Go was considered impossible for AI, and how deep learning plus Monte Carlo tree search produced AlphaGo's historic 2016 win over Lee Sedol. He details the leap to AlphaGo Zero and AlphaZero, which learned entirely from self-play with no human data, and MuZero, which learns even without being told the rules. The conversation closes on creativity, intrinsic reward, and a layered view of the meaning of life and intelligence.
Jonathan Williams
“let me mention that cryptocurrency in the context of the history of money it's fascinating I recommend a cent of money as a great book on this history”— Lex Fridman
Richard Sutton and Andrew Barto
“one of the things I read was Saturn Umberto the sort of seminal textbook an introduction to reinforcement learning and when I read that textbook I just had this resonating feeling”— guest
Stanley Kubrick
“almost like a fearful aw you know it's like in space 2001 Space Odyssey kind of realizing that you've created something”— Lex Fridman
Acorn Computers
“my my parents brought home this BBC modeled B microcomputer it was just this fascinating thing to me I was about seven years old”— guest
Will Wright
“will write the creator of SimCity and Sims on game design jane goodall on conservation Carlos Santana on guitar”— Lex Fridman
Will Wright
“will write the creator of SimCity and Sims on game design jane goodall on conservation Carlos Santana on guitar”— Lex Fridman