Home Lex Fridman Episode
Lex Fridman · 2022-12-06

Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation | Lex Fridman Podcast #344

AI researcher Noam Brown explains how his bots conquered poker and the negotiation game Diplomacy, and what that reveals about trust, search, and intelligence.

Noam Brown: AI vs Humans in Poker and Games of Strategic Negotiation | Lex Fridman Podcast #344
The guest

Noam Brown: Research scientist at Meta AI (FAIR) who co-created the first superhuman poker AIs (Libratus and Pluribus) and Cicero, an AI that negotiates with humans in natural language to play the board game Diplomacy at a human level.

What this episode covers

Noam Brown walks Lex Fridman through his career building game-playing AI, starting with No Limit Texas Hold'em. He explains Nash equilibrium, counterfactual regret minimization, and why real-time search dramatically boosts performance. He recounts the 2017 Libratus competition where his bot beat four top pros out of $2M over 120,000 hands, and how Pluribus extended this to six-player poker for a fraction of the cost. The conversation then turns to Diplomacy, a seven-player negotiation game that demands cooperation and natural language, requiring human data rather than pure self-play. Brown closes on trust, lying, cheat detection, data efficiency, and what it might take to reach AGI.

Recommended on this episode

MediaRecommended

Skyrim

Bethesda Game Studios

“what I think is the greatest game of all time”
“the Creator what I think is the greatest game of all time which is Skyrim and the NPCs there the AI that governs that whole game is very interesting”— Lex Fridman

Also referenced (named, not recommended)

MediaReferenced

Casino Royale

Martin Campbell

“if you watch movies about poker like Casino Royale or rounders the game that they're playing is no limit Texas hold in poker”— Noam Brown
MediaReferenced

Rounders

John Dahl

“if you watch movies about poker like Casino Royale or rounders the game that they're playing is no limit Texas hold in poker”— Noam Brown
MediaReferenced

Civilization

Sid Meier / Firaxis

“you look at a game like Civilization the way that the AIS play is not optimal for trying to win they're playing a different game”— Noam Brown
MediaReferenced

Starfield

Bethesda Game Studios

“Starfield the new game coming out and the Creator what I think is the greatest game of all time which is Skyrim”— Lex Fridman
MediaReferenced

The Elder Scrolls VI

Bethesda Game Studios

“Elder Scrolls 6 is in development now they're probably like pretty close to finishing it but I would not be surprised at all if Elder Scrolls 7 was using large language models”— Noam Brown
MediaReferenced

Diplomacy

Avalon Hill

“diplomacy has a much bigger Cooperative element to it it's a seven player game it was actually created in the 50s”— Lex Fridman
MediaReferenced

Settlers of Catan

Klaus Teuber

“you could have a natural language Settlers of Catan AI but the things that you're going to talk about are basically like am I trading you two sheep for a wood”— Noam Brown
MediaReferenced

Game of Thrones

HBO

“it's kind of like you know Game of Thrones like you got the the others coming from the north and like all the people have to start work out their differences”— Noam Brown

Big reveals from this episode

  • Libratus beat four top heads-up pros over 120,000 hands, winning close to $2 million.
  • The bot's surprising last-minute 'overbets' (betting many times the pot) flummoxed pros and have since become standard high-level poker strategy.
  • Without test-time search, the strongest Go bot drops from ~5200 ELO to ~3000 ELO; no raw neural net alone is superhuman at Go.
  • Brown admits the Libratus competition was so stressful he worked on it nearly nonstop for a year, and only gave the team ~50/50 odds going in.
  • Pluribus's final training run cost under $150 on AWS versus an estimated ~$100,000 for Libratus, driven by algorithmic gains (depth-limited search).
  • Libratus and Pluribus used no neural networks at all, which surprised many in the field.
  • Cicero was deliberately built to minimize lying, because lying made the bot perform worse once other players stopped trusting it.
  • A self-play bot trained without human data got destroyed by humans at Diplomacy even in the no-language version, because it couldn't model human conventions.

Worth remembering

  • Heads-up No Limit Texas Hold'em has about 10^161 decision points, more than the number of atoms in the universe squared.
  • In any finite two-player zero-sum game there is an optimal strategy that guarantees you won't lose in expectation, no matter what your opponent does.
  • The team handed the human pros the bot's hole cards for every hand each night, golden information, and still won.
  • Diplomacy was reportedly a favorite game of JFK and Henry Kissinger and was created in the 1950s, set just before World War I.
  • In the game's ideal form nobody wins, because if someone nears victory the others should unite to stop them, mirroring the futility of war.
  • Cicero's training data came from webDiplomacy.net: ~50,000 games and over 10 million natural-language messages, now being made available to researchers.
  • Cicero placed second out of ~19 players who played five or more games across 40 online Diplomacy games.
  • Brown argues Diplomacy is a bigger step toward real-world AI than StarCraft or Dota because it runs on open-ended natural language.
  • Humans show an 'anti-AI bias': told a bot was present, they spent whole games hunting it and ganging up to eliminate it.
  • Data inefficiency is a core gap: a Go AI needs millions of games to learn what a human grandmaster grasps in thousands.