Home Lex Fridman Episode
Lex Fridman · 2020-04-03

David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning | Lex Fridman Podcast #86

DeepMind's David Silver explains how AlphaGo, AlphaZero, and MuZero used self-play reinforcement learning to master games and discover superhuman creativity.

David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning | Lex Fridman Podcast #86
The guest

David Silver: Leader of DeepMind's reinforcement learning research group and the lead researcher on AlphaGo and AlphaZero, who also co-led AlphaStar and MuZero. He is one of the central figures behind modern deep reinforcement learning.

What this episode covers

David Silver traces his path from programming a BBC Micro at age seven and building games to a PhD applying reinforcement learning to the game of Go. He explains the core of reinforcement learning, why Go was considered impossible for AI, and how deep learning plus Monte Carlo tree search produced AlphaGo's historic 2016 win over Lee Sedol. He details the leap to AlphaGo Zero and AlphaZero, which learned entirely from self-play with no human data, and MuZero, which learns even without being told the rules. The conversation closes on creativity, intrinsic reward, and a layered view of the meaning of life and intelligence.

Recommended on this episode

BookRecommended

A Cent of Money: A History of Money

Jonathan Williams

“I recommend a cent of money as a great book on this history”
“let me mention that cryptocurrency in the context of the history of money it's fascinating I recommend a cent of money as a great book on this history”— Lex Fridman

Also referenced (named, not recommended)

BookReferenced

Reinforcement Learning: An Introduction

Richard Sutton and Andrew Barto

“one of the things I read was Saturn Umberto the sort of seminal textbook an introduction to reinforcement learning and when I read that textbook I just had this resonating feeling”— guest
MediaReferenced

2001: A Space Odyssey

Stanley Kubrick

“almost like a fearful aw you know it's like in space 2001 Space Odyssey kind of realizing that you've created something”— Lex Fridman
ProductReferenced

BBC Micro Model B

Acorn Computers

“my my parents brought home this BBC modeled B microcomputer it was just this fascinating thing to me I was about seven years old”— guest
MediaReferenced

SimCity

Will Wright

“will write the creator of SimCity and Sims on game design jane goodall on conservation Carlos Santana on guitar”— Lex Fridman
MediaReferenced

The Sims

Will Wright

“will write the creator of SimCity and Sims on game design jane goodall on conservation Carlos Santana on guitar”— Lex Fridman

Big reveals from this episode

  • A pure deep learning AlphaGo system with no search at all reached master-level Go, a definitive break from decades of search-dominated AI.
  • Silver predicted Lee Sedol 4-1 based on data showing AlphaGo developed a 'delusion' roughly 1 in 5 games.
  • He admits AlphaGo had inner 'holes' in its knowledge that persisted for tens of moves, and Lee Sedol exploited exactly that in the one game he won.
  • AlphaGo's 'move 37' broke every convention Go players are taught, proving machines could exhibit genuine creativity.
  • The full AlphaZero algorithm came to Silver while on his honeymoon, in his most relaxed state.
  • Silver offers a falsifiable prediction that scaling AlphaZero would beat each prior version 100-0 indefinitely throughout his lifetime.
  • MuZero learns a model of Atari, Go, chess and shogi without ever being told the rules, then plans to superhuman level.
  • World champion Magnus Carlsen credits studying AlphaZero's games for a new peak in his rating.

Worth remembering

  • In the 90s, heuristic search beat human world champions at chess, checkers, backgammon and Othello, but Go resisted.
  • When a $1M Go prize expired in 2000, the strongest program lost to a nine-year-old child even with a nine-stone handicap.
  • Go has the same number of stones for both players, so position value rests almost entirely on intuition, unlike chess's point system.
  • Go has around 10^170 possible positions, more than the roughly 10^80 atoms in the universe.
  • Monte Carlo search evaluates a position by playing random games to the end and averaging who wins.
  • After Lee Sedol, the next AlphaGo beat other top human players 60 games to nil.
  • AlphaZero's algorithm was independently applied in Nature papers to chemical synthesis and quantum computation, beating state of the art.
  • Silver borrows from Max Tegmark a view that the universe's 'goal' may be to maximize entropy, framing evolution and intelligence as nested sub-goals.
  • AlphaGo Zero rediscovered human joseki opening patterns then invented new variations now studied in top human competitions.