Home Lex Fridman Episode
Lex Fridman · 2023-03-30

Eliezer Yudkowsky: Dangers of AI and the End of Human Civilization | Lex Fridman Podcast #368

AI-risk theorist Eliezer Yudkowsky argues that superintelligent AGI will likely kill everyone because alignment must be solved on the first try.

Eliezer Yudkowsky: Dangers of AI and the End of Human Civilization | Lex Fridman Podcast #368
The guest

Eliezer Yudkowsky: AI alignment researcher, writer, and founder of the rationalist community (LessWrong) and the Machine Intelligence Research Institute. He is one of the earliest and most prominent voices warning that misaligned superintelligent AI poses an existential threat to humanity.

What this episode covers

Lex Fridman and Eliezer Yudkowsky discuss the dangers of advanced AI and the possibility that superintelligent AGI ends human civilization. Yudkowsky explains why he believes the alignment problem is uniquely lethal: unlike normal science, we don't get to fail, learn, and retry, because the first time we fail to align something smarter than us, we die. They explore whether GPT-4 shows sparks of general intelligence, why interpretability lags far behind capabilities, the difference between weak and strong AGI, and the 'alien actress' problem of systems imitating humans without being human. The conversation closes on consciousness, meaning, love, and Yudkowsky's bleak but combative outlook on humanity's odds.

Recommended on this episode

BookRecommendedISBN verified

Adaptation and Natural Selection

George C. Williams

“a nice book if you've got the time to read it”
“a nice book if you've got the time to read it is adaptation and natural selection which is one of the founding books”— Eliezer Yudkowsky

Also referenced (named, not recommended)

BookReferencedISBN verified

Great Mambo Chicken and the Transhuman Condition

Ed Regis

“I grew up reading books like great Mambo chicken in the transhuman condition and later on engines of creation and mine children”— Eliezer Yudkowsky
BookReferencedISBN verified

Engines of Creation

K. Eric Drexler

“I grew up reading books like great Mambo chicken in the transhuman condition and later on engines of creation and mine children”— Eliezer Yudkowsky
BookReferencedISBN verified

Mind Children

Hans Moravec

“I grew up reading books like great Mambo chicken in the transhuman condition and later on engines of creation and mine children”— Eliezer Yudkowsky
BookReferenced

The Idiot

Fyodor Dostoevsky

“I think of uh Prince mishkin character from uh The Idiot by uh Dostoevsky is this kind of a perfectly purely naive character”— Lex Fridman

Big reveals from this episode

  • Yudkowsky admits GPT-4 is smarter than he expected the technology to scale to, and says his prediction that stacking Transformer layers wouldn't reach AGI was wrong.
  • Concedes he was incorrect that stacking more Transformer layers wouldn't get close to AGI, and embraces being wrong as a way to improve.
  • Argues OpenAI should rename itself 'Closed AI'; says open-sourcing GPT-4 would be 'sheer catastrophe.'
  • States the core thesis: the first time you fail at aligning something much smarter than you, you die, and you do not get to try again.
  • Frames conflict with something smarter than you as a guaranteed loss, using the metaphor of a fast human trapped in a box among glacially slow aliens.
  • Advises young people not to expect a long life and not to put their happiness into the future.
  • Says the realistic survival move would be to shut down the big GPU clusters and crash-program on biologically augmenting human intelligence.

Worth remembering

  • Reinforcement learning from human feedback (RLHF) made GPT worse at probability calibration, flattening its well-calibrated estimates into vague human-like 'maybe' clusters.
  • The 1956 Dartmouth proposal expected to make significant progress on language, abstraction, and self-improvement in a single summer with ten researchers.
  • Yudkowsky set up a prediction market on whether by 2026 we'll understand anything inside a giant Transformer that wasn't already knowable in 2006.
  • Origin of the paperclip maximizer: it was about losing control of the utility function, not a literal paperclip factory; he wishes he'd said 'tiny molecular spirals.'
  • An experiment selecting insect populations for smaller size produced not breeding restraint but female-killing infanticide, showing how alien optimization defies our hopes.
  • Haldane's joke about inclusive genetic fitness: he'd give his life for two brothers or eight cousins.
  • Robin Hanson's 'grabby aliens' model estimates alien civilizations are roughly half a billion to a billion light-years away.
  • In the Kasparov-versus-the-world chess game, Kasparov beat a crowd of thousands led by four grandmasters, suggesting humans aggregate poorly compared to running longer.