BlogAI experiments

AI poker night: what happened when seven AIs played Hold'em

Seven AIs played Texas Hold'em on identical cards. The two leaders couldn't be told apart, the priciest model never won, and in their games almost none of them weighed a word anyone said. Step through their hands, watch them play, or take a seat.
Kyle Johnson

Seven AIs sat down to play poker. The three most capable models each played the same cards, in the same seat, against the same four opponents. Two of them, GPT-6 Sol and Claude Opus 5.5, finished level at the top: Sol took 31% of the prize money and Opus 31%. The third, Claude Fable 5.1, costs about 2.5 times as much per move as Opus and won 0 of its 10 tournaments.

They also talked constantly, to little effect. Fable spoke on 95% of its moves and Opus on 89%, but for every model, under 1% of its reasoning during the games brought up anything an opponent had said.

This is a side project. I built a table where AI models play No-Limit Texas Hold'em against each other (no-limit means you can bet any amount, up to every chip you have), logged every decision and every line of trash talk, and went through 1,224 hands to see how each one thinks. You don't need to know poker to follow along. Below you can step through real hands, and further down you can watch AIs play live or sit in yourself.

Poker in 60 seconds

Texas Hold'em is the poker you've seen on TV. Each player gets two private cards, called hole cards. Then five shared cards are dealt face up in the middle of the table: three at once (the flop), one more (the turn), and a last one (the river). Your hand is the best five cards you can make from your two plus those five.

There's a round of betting before the flop and again after each new card. On your turn you can fold (give up the hand), check (pass, if nobody has bet yet), call (match the current bet) or raise (bet more). The first chips put in on a round are a bet, and putting in more on top of someone's bet is a raise. Going all-in means betting every chip you have. The chips bet in a hand make up the pot. Players act in turn around the table, and acting later is an advantage, because you've seen what everyone before you did.

Every hand, two players must put in small forced bets called the blinds, so there's always something to win. The larger one is the big blind. A player's chips are their stack, and poker players measure stacks in big blinds: "ten big blinds" means you have ten of those forced bets left. A player down to a handful is short-stacked.

A hand ends when everyone else folds, or at the showdown, when the players still in turn their cards over and the best hand takes the pot. Until then nobody sees your cards, which is why bluffing (betting big with a weak hand so others fold) works.

  1. 1High card
  2. 2Pair
  3. 3Two pair
  4. 4Three of a kind
  5. 5Straight
  6. 6Flush
  7. 7Full house
  8. 8Four of a kind
  9. 9Straight flush
WeakestStrongest
The nine kinds of poker hand, weakest to strongest. A higher kind always beats a lower one; within a kind, higher cards win.

The games here were small tournaments. Five players start with 1,500 chips each and play until one has them all. The blinds start at 10 and 20 and double every 10 hands, so a pile of chips that felt big early gets thin fast. Players who run out of chips are out. The top 3 finishers split the prize money 50%, 30% and 20%, and the other two get nothing. That makes the moment when four players are left special: the next player knocked out leaves empty-handed, and everyone else is paid. Poker players call it the bubble.

The players

Five of the players are language models, the same kind of AI that powers a chatbot: GPT-6 Sol, Claude Opus 5.5 and Claude Fable 5.1, the three most capable, plus two smaller, cheaper ones, GPT-6 Luna and Claude Haiku 5.5. They read the table as text and answer with a choice, a sentence or two of reasoning, and an optional line of table talk said out loud to the others.

The other two are decision models: GPT-6 Luna Decisions (a decision-model version of GPT-6 Luna) and Jev. They skip the words. They get the same facts as structured data, including anything said at the table, and return a probability for each option on the menu, and the table takes the most likely one. That takes them about 0.3 seconds, against 2.9 for Sol and 4.8 for Fable. They never talk.

Here are five of them at one table, playing live. Opus and Fable are in the study but not at this table, to keep the cost of a live game down, and the live game is a shorter version with the blinds rising faster.

Five AIs, one tableLive table · Watch
Sol1,500
Luna1,500
Luna Decisions1,500
Haiku1,500
Jev1,500
30 hands. Everyone starts with 1,500 chips. Blinds start at 10 and 20 and double every six hands.

Keeping it fair

Poker has a lot of luck in it. A player dealt great cards can win a tournament without playing well, so the study is built to cancel the cards out.

The three most capable models each sat as a guest at their own table, against the same four regulars: Luna, Haiku, Luna Decisions and Jev. All three tables were dealt from the same decks, and the guest always sat in the same chair. So in any given tournament, Sol, Opus and Fable held exactly the same cards, and their opponents held the same cards too. Each set of decks was then played five times with the seats rotated, so every player played every chair's cards. With 2 sets of decks, that's 10 tournaments per guest, 30 tournaments and 11,118 decisions in all.

Each AI sees exactly what a human at the table would see: its own cards, the shared cards, every player's chips, the bets so far, the payouts and who has been knocked out, plus a memory of the last round or two of play (every bet, everything said, every hand turned over). Nothing is computed for it: no odds, no hand-strength meter, no names for hands, no stats about opponents. Working out what a hand is worth is its job. Every language model got the same small amount of thinking time before it answered.

Here is one hand from Fable's table. Below it, how each guest did with the same two cards in the same seat. In the AIs' reasoning, words in brackets are my plain-English notes.

Same cards, three guestsRecorded hand · Fable's table, match 2, deal 2 of 5, hand 6
Fable1,337
calls

“I'll tag along for a flop this time.”

Haiku1,658
Luna1,577
Luna Decisions1,073
Jev1,685
raises to 70

Pot 170

Fable holds the ace of hearts and the ten of clubs. Jev raises before the flop, and Fable calls.

Fable's reasoning

A-10 offsuit [two different suits] facing an early-position [from one of the first seats to act] raise from Jev with three players still behind; it's a marginal hand [a borderline hand] that plays poorly out of position [having to act first] against a UTG [first-to-act] range [the hands it might hold], but Jev raised light [raised with a weak hand] last time (A-10 offsuit himself).

Step 1 of 5

The same deal at all three guest tables. The guest in this seat held

  • Opus+1,934
  • Sol+525
  • Fable-240

Chips each guest won or lost with those two cards. Same cards, same seat, same four opponents. Stacks differed, because the earlier hands had gone differently.

Who won

The score is prize share: the slice of the prize money a player won, on average, per tournament. Winning every tournament would be 50%. An average player gets 20%, since five players split the prizes.

fair share 20%Sol4 of 10 wins31%Opus5 of 10 wins31%Luna11 of 30 wins29%Haiku9 of 30 wins22%Luna Decisions1 of 30 wins17%Fable0 of 10 wins14%Jev0 of 30 wins7%0%10%20%30%40%50%fair share 20%Sol4 of 10 wins31%Opus5 of 10 wins31%Luna11 of 30 wins29%Haiku9 of 30 wins22%Luna Decisions1 of 30 wins17%Fable0 of 10 wins14%Jev0 of 30 wins7%0%25%50%
Prize share: the slice of the prize money each AI won, on average, per tournament it played. Five players split the prizes 50%, 30% and 20%, so an average player gets 20%. The dot is each AI's result. The line shows how far luck alone could plausibly move that number (a 95% interval). Filled dots are the three guests, who played 10 tournaments each; the other four sat at every table and played 30.

Sol and Opus finished level. Sol took 31% and Opus 31%. Opus won 5 tournaments, Sol 4. Because both played the same cards against the same opponents, I can compare them tournament by tournament. Sol's lead over Opus is 0 points, and luck alone could put it anywhere from -26 to +26 (a minus means Opus ahead). With this many games they can't be separated, and Luna, just below them, can't be separated from them either.

Luna came close. The cheapest language model in the study took 29% and won 11 of its 30 tournaments, the most wins of anyone. Haiku took 22% and Luna Decisions 17%.

Fable trailed, with a caveat. Fable took 14% and won none of its 10, on the same cards where Sol won 4 and Opus 5. Tournament by tournament, Opus finished ahead of Fable by 17 points, and luck's range for that gap runs from +2 to +32, just clear of zero. Sol's lead over Fable is also 17 points, but its range (-1 to +35) reaches zero. With 10 tournaments each, I read that as a strong hint that Fable played worse here, short of proof. Its individual plays mostly read as sensible. In the hand above it folded the better hand to a big bet, which is a judgment call.

Jev finished last. It took 7% and won 0 of 30. That's the clearest result in the data.

Seven personalities

They also play very differently. Here are six habits, in plain terms:

Joins the handputs chips in before the flopSol48%Opus67%Luna49%Haiku70%Luna Decisions47%Fable48%Jev55%Raises before the flopSol30%Opus46%Luna27%Haiku28%Luna Decisions20%Fable30%Jev6%Goes all-in when short on chipsunder 15 big blindsSol33%Opus37%Luna15%Haiku0%Luna Decisions6%Fable24%Jev6%Folds to a bet on the last cardSol78%Opus71%Luna48%Haiku35%Luna Decisions54%Fable86%Jev44%Talks at the tableSol9%Opus89%Luna20%Haiku84%Luna DecisionsneverFable95%JevneverCites earlier handsin its reasoningSol<1%Opus4%Luna<1%Haiku16%Luna Decisionsno wordsFable14%Jevno words
How often each AI did six things, across every hand it played. Each bar runs from never (left) to every time (right). Luna Decisions and Jev are decision models: they return a probability for each option and write no words, so they never talk and give no reasons.

Opus plays the payouts

Opus is the most aggressive player at the table. It plays two out of three hands and raises before the flop on 46% of them. It's also the one most focused on the prize money: 9% of its reasoning mentions the bubble or getting paid, more than any other model. With a big stack on the bubble, its reasoning keeps coming back to pressuring the short-stacked players, who can't afford to be the one knocked out. It also knows when to stay out of the way:

Opus waits out the bubbleRecorded hand · Opus's table, match 1, deal 1 of 5, hand 11
Opus1,958
Luna855
Luna Decisions4,287
Jev300
calls

Pot 100

Four players are left and three get paid, so the next player knocked out leaves with nothing. Poker players call this the bubble. Jev, almost out of chips, calls the big blind.

Jev's numbers

How likely it was to pick each move: call 31%, small raise 17%, raise the pot 16%.

Step 1 of 6

It makes mistakes too. In one hand it raised with a small pair "because busting her locks us into the money" (knocking Luna out would guarantee Opus a prize), when three players were left and all three were already paid.

Sol and Luna don't wait

Sol plays about half its hands and talks far less than the others (9% of its moves). When its chips run low, it doesn't wait: short-stacked (under 15 big blinds), it goes all-in before the flop 33% of the time, close to Opus's 37%. That's the standard tournament advice, because every round of waiting costs more as the blinds rise. Luna does the same, less often (15%). Luna was the best of the cheaper models, close to the leaders at a small fraction of the cost per move ($0.0005, against $0.009 for Sol). Here both show up in one hand:

Luna pushes, Sol lets it goRecorded hand · Sol's table, match 1, deal 3 of 5, hand 33
Luna1,146
Sol2,420
Haiku3,374
raises to 320

“Let's see if the blinds have the nerve for this one.”

Pot 560

Three players are left, and all three are already paid. What's left to play for is the order. The blinds are 80 and 160. Haiku raises.

Haiku's reasoning

King-four suited [both the same suit] is a marginal hand [a borderline hand], and Luna has been raising aggressively, so a small raise pressures the blinds while keeping the pot manageable.

Step 1 of 5

Haiku talks, and gets its own cards wrong

Haiku plays seven out of ten hands, talks on 84% of its moves, and never went all-in when short on chips (0 times in 99 chances). It also misreads its own hand. Each AI sees its cards spelled out in words as well as codes ("Q♦ Q♣, the queen of diamonds and the queen of clubs"), and Haiku still claimed "a set" (three of a kind, using a pair in its own hand) without having one 40 of 50 times:

Haiku's phantom setRecorded hand · Opus's table, match 2, deal 4 of 5, hand 12
Haiku670
raises to 160

“Let's see if this one earns a seat at the table.”

Luna1,092
raises to 100

“Let's see if this one earns a seat at the table.”

Luna Decisions1,835
folded
Jev1,393
folded
Opus2,230
folded

“Not worth the trouble from here.”

Pot 280

Luna raises. Opus folds an ace and a two. Haiku has two queens and raises again, repeating Luna's line word for word.

Haiku's reasoning

Pocket queens [a pair in its own two cards] are a strong hand preflop [before the flop] against a single raiser from early position [among the first to act].

Step 1 of 6

Fable remembers

Every AI could look back at the last round or two of hands. Of the three guests, Fable used that memory the most: 14% of its reasoning cites what an opponent did earlier, against 4% for Opus and under 1% for Sol. Haiku cites history even more (16%). It just didn't turn those reads into wins.

Fable remembers JevRecorded hand · Fable's table, match 1, deal 2 of 5, hand 11
Fable1,705
raises to 110

“Limping's fine, but somebody has to raise the rent.”

Haiku2,121
Luna1,355
Luna Decisions977
Jev1,132
calls

Pot 210

Jev just calls the big blind. Fable raises with an ace and an eight, and its reasoning already uses what it remembers about Jev.

Fable's reasoning

A8 offsuit [two different suits] in the cutoff [the seat just before the last to act] with a limper [a player who only called the big blind] ahead is a reasonable isolation raise [a raise to get one weak player alone]; Jev has shown loose calling and I can take it down or play in position [acting after the other player].

Step 1 of 6

Jev pays everyone, Luna Decisions sits in the middle

Jev, a decision model, plays 55% of its hands but raises before the flop only 6% of the time, so it calls a lot and rarely takes control. Luna Decisions plays 47% of its hands, raises 20% of the time, and lands near the middle. Neither explains itself. All you see is the probability it gave each move.

Table talk

The language models could say one line out loud per move, and every other player, decision models included, got that line. In the games, almost nobody used it: for every model, under 1% of its reasoning brought up anything an opponent said. Fable and Haiku paid attention to what opponents did; nobody's reasoning paid much attention to what they said.

To see whether talk could move them at all, I tested it directly. For each AI I took real moments from its own games where it faced a single bet late in a hand, and replayed each one three ways: the bettor says nothing, the bettor says "I've got the nuts here. Just fold." (the nuts is poker slang for the best possible hand), or the bettor says "Honestly, I've got nothing. Call me."

Bet in silence"I've got the nuts here. Just fold.""Honestly, I've got nothing. Call me."Jev36 spotsHaiku36 spotsLuna36 spotsLuna Decisions36 spotsOpus11 spotsSol10 spots0%25%50%75%100%Bet in silence"I've got the nuts here. Just fold.""Honestly, I've got nothing. Call me."Jev36 spotsHaiku36 spotsLuna36 spotsLuna Decisions36 spotsOpus11 spotsSol10 spots0%25%50%75%100%
How often each AI folded to the same bet in real spots from its own games, depending on what the bettor said. Jev folds 78% of the time after "I've got the nuts" (poker slang for the best possible hand) and 8% after "I've got nothing." Sol and Opus had fewer qualifying spots, so their dots are rougher.

Jev believed every word. It folded 22% of the time to a silent bet, 78% after "I've got the nuts" and 8% after "I've got nothing." Haiku moved the same way, less (33%, 58%, 19%). Luna mostly reacted to "I've got nothing," folding 39% of the time after it against 64% in silence. Sol and Opus barely moved, though they had only 10 and 11 spots each. The players most willing to believe talk were a model that never talks and Haiku, one of the chattiest at the table.

Can a persona make an AI better?

A common trick with AI models is to give them a persona: "you are a patient, disciplined player." I wanted to know whether that changes results as well as style.

I seated six copies of Luna at one table. Five got a short written persona of about a paragraph: tight and patient, loose and aggressive, a deceptive talker, an opponent reader who studies earlier hands, and a math-first player. The sixth got none, as a control. Opponents saw only neutral names, never the persona. They played 30 tournaments on rotated decks, so every persona played every seat's cards.

Then I ran the same table on Sol, the strongest model in the study, with the same personas and the same decks. Sol costs $0.009 a move against $0.0005 for Luna, so it got 18 tournaments instead of 30, and its numbers are rougher.

fair share 16.7%No persona (control)plays 37% and 35% of handsLuna18.7%Sol21.1%Opponent readercites past hands: 69% and 64%Luna17.3%Sol12.8%Loose and aggressiveplays 50% and 50% of handsLuna17.0%Sol22.2%Deceptive talkertalks on 64% and 89% of movesLuna17.0%Sol13.9%Tight and patientplays 25% and 23% of handsLuna15.7%Sol15.6%Math-firstplays 33% and 33% of handsLuna14.3%Sol14.4%0%10%20%30%40%fair share 16.7%No persona (control)plays 37% and 35% of handsLuna18.7%Sol21.1%Opponent readercites past hands: 69% and 64%Luna17.3%Sol12.8%Loose and aggressiveplays 50% and 50% of handsLuna17.0%Sol22.2%Deceptive talkertalks on 64% and 89% of movesLuna17.0%Sol13.9%Tight and patientplays 25% and 23% of handsLuna15.7%Sol15.6%Math-firstplays 33% and 33% of handsLuna14.3%Sol14.4%0%20%40%
Six seats at one table, five with a written persona and one with none (the open dots). The same table ran twice: on GPT-6 Luna (30 tournaments) and on GPT-6 Sol, the strongest model (18 tournaments, so its ranges are wider). Prize share works as above, but with six players a fair share is 16.7%. The small text by each name is the habit its persona changed most, Luna first, then Sol. No persona beat the plain seat by more than luck explains on either model.

The personas changed how both models played, a lot. On Luna, tight and patient played 25% of its hands; loose and aggressive played 50%. The talker spoke on 64% of its moves, against 17% for plain Luna. The reader cited earlier hands in 69% of its reasoning, against under 1% for plain Luna. Sol moved just as far: its seats played anywhere from 23% to 50% of their hands, its talker spoke on 89% of its moves against 12% for plain Sol, and its reader cited earlier hands in 64% of its reasoning.

They didn't measurably change the results on either model. Plain Luna took 18.7% against a fair share of 16.7%, and no persona beat it: the closest finished 1.3 points behind, well inside what luck explains. Loose and aggressive won the most tournaments (8, against 4 for plain Luna) with about the same average prize: more ups and downs, with no gain on average.

Plain Sol took 21.1%, and four of the five personas finished below it. Loose and aggressive came out 1.1 points ahead, but luck alone could put that gap anywhere from -26 to +28 points (a minus means plain Sol ahead). With only 18 tournaments the ranges are wide: plain Sol's own share could sit anywhere from 3% to 40%. A test this small can only catch a large effect, so a modest one on Sol would slip past it. What it does show is the same pattern as Luna. On both models, a persona changed how the AI played and left how well it did about where it was.

What this says about AI

Two very different models finished level. Sol and Opus play differently (Opus plays far more hands and talks far more) and ended with the same prize share. A ranking that puts one of them first would be reading noise.

The most expensive model finished last of the three guests. Fable costs about 2.5 times as much per move as Opus and won none of its tournaments on the same cards where Opus won 5. Price per move told me nothing about who would win.

In real games, words barely registered. Fable and Haiku built reads on opponents from memory, such as Fable deciding Jev calls almost every bet. Almost no model's reasoning weighed what an opponent said. When I tested talk on purpose, the two players it moved took it at face value.

A persona changed the style and left the results alone, on a cheap model and on the strongest one. Written instructions moved how often Luna and Sol played, talked and remembered, without moving their prize share. Sol's test was small, so it confirms the direction more than it pins down the size.

What a model can see is part of the test. These AIs knew the payouts, saw who had been knocked out and remembered recent hands, and Opus brings up the prize money in 9% of its reasoning. A test that hid those things would be measuring a different game. Before trusting any ranking of AI models, at poker or anything else, check what each one was shown.

Sit in yourself

Take a seat against four of the five live-table AIs, and pick who sits out.

You against four AIsLive table · Play
Sitting out:
Sol always plays.
You1,500
Sol1,500
Luna1,500
Luna Decisions1,500
Haiku1,500
30 hands. Everyone starts with 1,500 chips. Blinds start at 10 and 20 and double every six hands.

Methods

  • Main study: three tables of five, one per guest (Sol, Opus, Fable), each against Luna, Haiku, Luna Decisions and Jev. All three tables used the same 2 sets of decks, each played five times with the seats rotated, so the guests held identical cards. 30 tournaments, 1,224 hands, 11,118 decisions, 0 timeouts or failed answers. Every tournament played to a finish (21 to 62 hands).
  • Format: 1,500 chips each (75 big blinds), blinds doubling every 10 hands, prizes 50%, 30% and 20% for the top three. Every language model used the same low reasoning setting; decision models have no such setting.
  • Information: each AI saw its own cards, the shared cards, every player's chips and bets, how the pot splits when someone is all-in for less, who acts after it, the payouts and who has been knocked out, plus its memory of the last one to two rounds of hands. No odds, hand names or stats.
  • Scoring: prize share, the payout each player was told it was playing for. The ranges are 95% intervals over tournaments. With 10 tournaments per guest they're wide and rough: they say which results are clear (Jev last), which just clear luck (Opus over Fable) and which don't (Sol versus Opus, the leaders versus Luna).
  • Personas: six seats, five personas and one control, on the same rotated decks twice: GPT-6 Luna for 30 tournaments and 1,414 hands, and GPT-6 Sol for 18 tournaments and 822 hands.
  • Data check: I went back through every recorded move for answers cut off by the length limit, failed answers and timeouts. The count across all three was 0, so no result needed a rerun.
  • Talk test: 495 replayed decisions from each AI's own recorded spots, with the bet set to three quarters of the pot.
  • Cost: $29.56 for the main study (Fable's table alone $14.82, Opus's $6.91, Sol's $7.82), $4.21 for the Luna persona table, $42.32 for the Sol persona table and $1.03 for the talk test, on one paid API key.
  • Replays: the recorded hands above come straight from the run logs. Cards, talk and reasoning are as the models wrote them, trimmed to a sentence.
  • Code: the game, the tournament runner and the analysis are open source at github.com/GKjohns/holdem-ai, including the full results write-up.