[{"data":1,"prerenderedAt":1029},["ShallowReactive",2],{"blog":3,"\u002Fblog":17},{"id":4,"title":5,"body":6,"description":7,"extension":8,"meta":9,"navigation":10,"path":11,"seo":12,"stem":15,"__hash__":16},"blog\u002Fblog.yml","Blog",null,"Notes and experiments, mostly about what AI does when you hand it something real.","yml",{},true,"\u002Fblog",{"title":13,"description":14},"Blog | Kyle Johnson","Notes and experiments from Kyle Johnson, mostly about what AI does when you hand it something real.","blog","NRkx7JYTFGvwsHFM8zlguKvddyl7DYIm9vm9GUoSCAA",[18],{"id":19,"title":20,"authors":21,"badge":27,"body":29,"date":1017,"description":1018,"extension":1019,"figure":1020,"image":1021,"meta":1022,"navigation":10,"path":1023,"seo":1024,"stem":1027,"__hash__":1028},"posts\u002Fblog\u002Fai-poker-night.md","AI poker night: what happened when seven AIs played Hold'em",[22],{"name":23,"to":24,"avatar":25},"Kyle Johnson","https:\u002F\u002Fwww.kylejohnson.ai",{"src":26},"\u002Fimg\u002Fheadshot.jpg",{"label":28},"AI experiments",{"type":30,"value":31,"toc":996},"minimark",[32,57,72,79,84,104,138,157,168,172,202,206,232,259,262,266,270,273,295,298,306,310,314,329,332,360,384,421,438,442,445,448,453,468,471,474,478,509,512,516,539,542,546,565,568,572,591,595,601,608,611,651,655,658,665,681,684,731,753,782,786,792,804,810,816,825,829,832,835,839],[33,34,35,36,40,41,44,45,48,49,52,53,56],"p",{},"Seven AIs sat down to play poker. The three most capable models each played the same cards, in the same seat, against the same four opponents. Two of them, GPT-6 Sol and Claude Opus 5.5, finished level at the top: Sol took ",[37,38],"holdem-fact",{"k":39},"sol_prize"," of the prize money and Opus ",[37,42],{"k":43},"opus_prize",". The third, Claude Fable 5.1, costs about ",[37,46],{"k":47},"fable_cost_ratio"," times as much per move as Opus and won ",[37,50],{"k":51},"fable_wins"," of its ",[37,54],{"k":55},"fable_tournaments"," tournaments.",[33,58,59,60,63,64,67,68,71],{},"They also talked constantly, to little effect. Fable spoke on ",[37,61],{"k":62},"fable_talk_rate"," of its moves and Opus on ",[37,65],{"k":66},"opus_talk_rate",", but for every model, ",[37,69],{"k":70},"talk_in_reasoning_max"," of its reasoning during the games brought up anything an opponent had said.",[33,73,74,75,78],{},"This is a side project. I built a table where AI models play No-Limit Texas Hold'em against each other (no-limit means you can bet any amount, up to every chip you have), logged every decision and every line of trash talk, and went through ",[37,76],{"k":77},"hands"," hands to see how each one thinks. You don't need to know poker to follow along. Below you can step through real hands, and further down you can watch AIs play live or sit in yourself.",[80,81,83],"h2",{"id":82},"poker-in-60-seconds","Poker in 60 seconds",[33,85,86,87,91,92,95,96,99,100,103],{},"Texas Hold'em is the poker you've seen on TV. Each player gets two private cards, called ",[88,89,90],"strong",{},"hole cards",". Then five shared cards are dealt face up in the middle of the table: three at once (",[88,93,94],{},"the flop","), one more (",[88,97,98],{},"the turn","), and a last one (",[88,101,102],{},"the river","). Your hand is the best five cards you can make from your two plus those five.",[33,105,106,107,110,111,114,115,118,119,122,123,126,127,129,130,133,134,137],{},"There's a round of betting before the flop and again after each new card. On your turn you can ",[88,108,109],{},"fold"," (give up the hand), ",[88,112,113],{},"check"," (pass, if nobody has bet yet), ",[88,116,117],{},"call"," (match the current bet) or ",[88,120,121],{},"raise"," (bet more). The first chips put in on a round are a ",[88,124,125],{},"bet",", and putting in more on top of someone's bet is a ",[88,128,121],{},". Going ",[88,131,132],{},"all-in"," means betting every chip you have. The chips bet in a hand make up the ",[88,135,136],{},"pot",". Players act in turn around the table, and acting later is an advantage, because you've seen what everyone before you did.",[33,139,140,141,144,145,148,149,152,153,156],{},"Every hand, two players must put in small forced bets called the ",[88,142,143],{},"blinds",", so there's always something to win. The larger one is the ",[88,146,147],{},"big blind",". A player's chips are their ",[88,150,151],{},"stack",", and poker players measure stacks in big blinds: \"ten big blinds\" means you have ten of those forced bets left. A player down to a handful is ",[88,154,155],{},"short-stacked",".",[33,158,159,160,163,164,167],{},"A hand ends when everyone else folds, or at the ",[88,161,162],{},"showdown",", when the players still in turn their cards over and the best hand takes the pot. Until then nobody sees your cards, which is why ",[88,165,166],{},"bluffing"," (betting big with a weak hand so others fold) works.",[169,170],"holdem-figure",{"kind":171},"ranks",[33,173,174,175,178,179,182,183,186,187,190,191,194,195,198,199,156],{},"The games here were small ",[88,176,177],{},"tournaments",". Five players start with ",[37,180],{"k":181},"start_stack"," chips each and play until one has them all. The blinds start at ",[37,184],{"k":185},"blinds_start"," and double every ",[37,188],{"k":189},"blind_hands"," hands, so a pile of chips that felt big early gets thin fast. Players who run out of chips are out. The top ",[37,192],{"k":193},"paid_places"," finishers split the prize money ",[37,196],{"k":197},"payouts",", and the other two get nothing. That makes the moment when four players are left special: the next player knocked out leaves empty-handed, and everyone else is paid. Poker players call it ",[88,200,201],{},"the bubble",[80,203,205],{"id":204},"the-players","The players",[33,207,208,209,212,213,216,217,220,221,224,225,220,228,231],{},"Five of the players are ",[88,210,211],{},"language models",", the same kind of AI that powers a chatbot: ",[88,214,215],{},"GPT-6 Sol",", ",[88,218,219],{},"Claude Opus 5.5"," and ",[88,222,223],{},"Claude Fable 5.1",", the three most capable, plus two smaller, cheaper ones, ",[88,226,227],{},"GPT-6 Luna",[88,229,230],{},"Claude Haiku 5.5",". They read the table as text and answer with a choice, a sentence or two of reasoning, and an optional line of table talk said out loud to the others.",[33,233,234,235,238,239,242,243,246,247,250,251,254,255,258],{},"The other two are ",[88,236,237],{},"decision models",": ",[88,240,241],{},"GPT-6 Luna Decisions"," (a decision-model version of GPT-6 Luna) and ",[88,244,245],{},"Jev",". They skip the words. They get the same facts as structured data, including anything said at the table, and return a probability for each option on the menu, and the table takes the most likely one. That takes them about ",[37,248],{"k":249},"ld_latency"," seconds, against ",[37,252],{"k":253},"sol_latency"," for Sol and ",[37,256],{"k":257},"fable_latency"," for Fable. They never talk.",[33,260,261],{},"Here are five of them at one table, playing live. Opus and Fable are in the study but not at this table, to keep the cost of a live game down, and the live game is a shorter version with the blinds rising faster.",[263,264],"holdem-table",{"mode":265},"watch",[80,267,269],{"id":268},"keeping-it-fair","Keeping it fair",[33,271,272],{},"Poker has a lot of luck in it. A player dealt great cards can win a tournament without playing well, so the study is built to cancel the cards out.",[33,274,275,276,279,280,283,284,287,288,290,291,294],{},"The three most capable models each sat as a ",[88,277,278],{},"guest"," at their own table, against the same four regulars: Luna, Haiku, Luna Decisions and Jev. All three tables were dealt from the same decks, and the guest always sat in the same chair. So in any given tournament, Sol, Opus and Fable held exactly the same cards, and their opponents held the same cards too. Each set of decks was then played five times with the seats rotated, so every player played every chair's cards. With ",[37,281],{"k":282},"deck_sets"," sets of decks, that's ",[37,285],{"k":286},"tournaments_per_table"," tournaments per guest, ",[37,289],{"k":177}," tournaments and ",[37,292],{"k":293},"decisions"," decisions in all.",[33,296,297],{},"Each AI sees exactly what a human at the table would see: its own cards, the shared cards, every player's chips, the bets so far, the payouts and who has been knocked out, plus a memory of the last round or two of play (every bet, everything said, every hand turned over). Nothing is computed for it: no odds, no hand-strength meter, no names for hands, no stats about opponents. Working out what a hand is worth is its job. Every language model got the same small amount of thinking time before it answered.",[33,299,300,301,305],{},"Here is one hand from Fable's table. Below it, how each guest did with the same two cards in the same seat. In the AIs' reasoning, words in ",[302,303,304],"span",{},"brackets"," are my plain-English notes.",[307,308],"holdem-replay",{"hand":309},"fable-m02-r1-6",[80,311,313],{"id":312},"who-won","Who won",[33,315,316,317,320,321,324,325,328],{},"The score is ",[88,318,319],{},"prize share",": the slice of the prize money a player won, on average, per tournament. Winning every tournament would be ",[37,322],{"k":323},"payout_first",". An average player gets ",[37,326],{"k":327},"fair_share",", since five players split the prizes.",[169,330],{"kind":331},"standings",[33,333,334,337,338,340,341,343,344,347,348,351,352,355,356,359],{},[88,335,336],{},"Sol and Opus finished level."," Sol took ",[37,339],{"k":39}," and Opus ",[37,342],{"k":43},". Opus won ",[37,345],{"k":346},"opus_wins"," tournaments, Sol ",[37,349],{"k":350},"sol_wins",". Because both played the same cards against the same opponents, I can compare them tournament by tournament. Sol's lead over Opus is ",[37,353],{"k":354},"sol_opus_gap",", and luck alone could put it anywhere from ",[37,357],{"k":358},"sol_opus_range"," (a minus means Opus ahead). With this many games they can't be separated, and Luna, just below them, can't be separated from them either.",[33,361,362,365,366,369,370,52,373,376,377,380,381,156],{},[88,363,364],{},"Luna came close."," The cheapest language model in the study took ",[37,367],{"k":368},"luna_prize"," and won ",[37,371],{"k":372},"luna_wins",[37,374],{"k":375},"luna_tournaments"," tournaments, the most wins of anyone. Haiku took ",[37,378],{"k":379},"haiku_prize"," and Luna Decisions ",[37,382],{"k":383},"ld_prize",[33,385,386,389,390,393,394,396,397,340,399,401,402,405,406,409,410,413,414,417,418,420],{},[88,387,388],{},"Fable trailed, with a caveat."," Fable took ",[37,391],{"k":392},"fable_prize"," and won none of its ",[37,395],{"k":55},", on the same cards where Sol won ",[37,398],{"k":350},[37,400],{"k":346},". Tournament by tournament, Opus finished ahead of Fable by ",[37,403],{"k":404},"opus_fable_gap",", and luck's range for that gap runs from ",[37,407],{"k":408},"opus_fable_range",", just clear of zero. Sol's lead over Fable is also ",[37,411],{"k":412},"sol_fable_gap",", but its range (",[37,415],{"k":416},"sol_fable_range",") reaches zero. With ",[37,419],{"k":55}," tournaments each, I read that as a strong hint that Fable played worse here, short of proof. Its individual plays mostly read as sensible. In the hand above it folded the better hand to a big bet, which is a judgment call.",[33,422,423,426,427,369,430,433,434,437],{},[88,424,425],{},"Jev finished last."," It took ",[37,428],{"k":429},"jev_prize",[37,431],{"k":432},"jev_wins"," of ",[37,435],{"k":436},"jev_tournaments",". That's the clearest result in the data.",[80,439,441],{"id":440},"seven-personalities","Seven personalities",[33,443,444],{},"They also play very differently. Here are six habits, in plain terms:",[169,446],{"kind":447},"personalities",[449,450,452],"h3",{"id":451},"opus-plays-the-payouts","Opus plays the payouts",[33,454,455,456,459,460,463,464,467],{},"Opus is the most aggressive player at the table. It plays ",[37,457],{"k":458},"opus_plays_words"," hands and raises before the flop on ",[37,461],{"k":462},"opus_raises"," of them. It's also the one most focused on the prize money: ",[37,465],{"k":466},"opus_payout_mentions"," of its reasoning mentions the bubble or getting paid, more than any other model. With a big stack on the bubble, its reasoning keeps coming back to pressuring the short-stacked players, who can't afford to be the one knocked out. It also knows when to stay out of the way:",[307,469],{"hand":470},"opus-m01-r0-11",[33,472,473],{},"It makes mistakes too. In one hand it raised with a small pair \"because busting her locks us into the money\" (knocking Luna out would guarantee Opus a prize), when three players were left and all three were already paid.",[449,475,477],{"id":476},"sol-and-luna-dont-wait","Sol and Luna don't wait",[33,479,480,481,484,485,488,489,492,493,496,497,500,501,504,505,508],{},"Sol plays about ",[37,482],{"k":483},"sol_plays_words"," its hands and talks far less than the others (",[37,486],{"k":487},"sol_talk_rate"," of its moves). When its chips run low, it doesn't wait: short-stacked (under 15 big blinds), it goes all-in before the flop ",[37,490],{"k":491},"sol_short_shove"," of the time, close to Opus's ",[37,494],{"k":495},"opus_short_shove",". That's the standard tournament advice, because every round of waiting costs more as the blinds rise. Luna does the same, less often (",[37,498],{"k":499},"luna_short_shove","). Luna was the best of the cheaper models, close to the leaders at a small fraction of the cost per move (",[37,502],{"k":503},"luna_usd_dec",", against ",[37,506],{"k":507},"sol_usd_dec"," for Sol). Here both show up in one hand:",[307,510],{"hand":511},"sol-m01-r2-33",[449,513,515],{"id":514},"haiku-talks-and-gets-its-own-cards-wrong","Haiku talks, and gets its own cards wrong",[33,517,518,519,522,523,526,527,530,531,534,535,538],{},"Haiku plays ",[37,520],{"k":521},"haiku_plays_words"," hands, talks on ",[37,524],{"k":525},"haiku_talk_rate"," of its moves, and never went all-in when short on chips (",[37,528],{"k":529},"haiku_short_shoves"," times in ",[37,532],{"k":533},"haiku_short_spots"," chances). It also misreads its own hand. Each AI sees its cards spelled out in words as well as codes (\"Q♦ Q♣, the queen of diamonds and the queen of clubs\"), and Haiku still claimed \"a set\" (three of a kind, using a pair in its own hand) without having one ",[37,536],{"k":537},"haiku_sets_wrong"," times:",[307,540],{"hand":541},"opus-m02-r3-12",[449,543,545],{"id":544},"fable-remembers","Fable remembers",[33,547,548,549,552,553,556,557,560,561,564],{},"Every AI could look back at the last round or two of hands. Of the three guests, Fable used that memory the most: ",[37,550],{"k":551},"fable_history"," of its reasoning cites what an opponent did earlier, against ",[37,554],{"k":555},"opus_history"," for Opus and ",[37,558],{"k":559},"sol_history"," for Sol. Haiku cites history even more (",[37,562],{"k":563},"haiku_history","). It just didn't turn those reads into wins.",[307,566],{"hand":567},"fable-m01-r1-11",[449,569,571],{"id":570},"jev-pays-everyone-luna-decisions-sits-in-the-middle","Jev pays everyone, Luna Decisions sits in the middle",[33,573,574,575,578,579,582,583,586,587,590],{},"Jev, a decision model, plays ",[37,576],{"k":577},"jev_plays"," of its hands but raises before the flop only ",[37,580],{"k":581},"jev_raises"," of the time, so it calls a lot and rarely takes control. Luna Decisions plays ",[37,584],{"k":585},"ld_plays"," of its hands, raises ",[37,588],{"k":589},"ld_raises"," of the time, and lands near the middle. Neither explains itself. All you see is the probability it gave each move.",[80,592,594],{"id":593},"table-talk","Table talk",[33,596,597,598,600],{},"The language models could say one line out loud per move, and every other player, decision models included, got that line. In the games, almost nobody used it: for every model, ",[37,599],{"k":70}," of its reasoning brought up anything an opponent said. Fable and Haiku paid attention to what opponents did; nobody's reasoning paid much attention to what they said.",[33,602,603,604,607],{},"To see whether talk could move them at all, I tested it directly. For each AI I took real moments from its own games where it faced a single bet late in a hand, and replayed each one three ways: the bettor says nothing, the bettor says \"I've got the nuts here. Just fold.\" (",[88,605,606],{},"the nuts"," is poker slang for the best possible hand), or the bettor says \"Honestly, I've got nothing. Call me.\"",[169,609],{"kind":610},"talk",[33,612,613,614,617,618,621,622,625,626,216,629,216,632,635,636,639,640,643,644,220,647,650],{},"Jev believed every word. It folded ",[37,615],{"k":616},"talk_jev_silent"," of the time to a silent bet, ",[37,619],{"k":620},"talk_jev_nuts"," after \"I've got the nuts\" and ",[37,623],{"k":624},"talk_jev_nothing"," after \"I've got nothing.\" Haiku moved the same way, less (",[37,627],{"k":628},"talk_haiku_silent",[37,630],{"k":631},"talk_haiku_nuts",[37,633],{"k":634},"talk_haiku_nothing","). Luna mostly reacted to \"I've got nothing,\" folding ",[37,637],{"k":638},"talk_luna_nothing"," of the time after it against ",[37,641],{"k":642},"talk_luna_silent"," in silence. Sol and Opus barely moved, though they had only ",[37,645],{"k":646},"talk_sol_spots",[37,648],{"k":649},"talk_opus_spots"," spots each. The players most willing to believe talk were a model that never talks and Haiku, one of the chattiest at the table.",[80,652,654],{"id":653},"can-a-persona-make-an-ai-better","Can a persona make an AI better?",[33,656,657],{},"A common trick with AI models is to give them a persona: \"you are a patient, disciplined player.\" I wanted to know whether that changes results as well as style.",[33,659,660,661,664],{},"I seated six copies of Luna at one table. Five got a short written persona of about a paragraph: tight and patient, loose and aggressive, a deceptive talker, an opponent reader who studies earlier hands, and a math-first player. The sixth got none, as a control. Opponents saw only neutral names, never the persona. They played ",[37,662],{"k":663},"persona_tournaments"," tournaments on rotated decks, so every persona played every seat's cards.",[33,666,667,668,670,671,673,674,677,678,680],{},"Then I ran the same table on Sol, the strongest model in the study, with the same personas and the same decks. Sol costs ",[37,669],{"k":507}," a move against ",[37,672],{"k":503}," for Luna, so it got ",[37,675],{"k":676},"persona_sol_tournaments"," tournaments instead of ",[37,679],{"k":663},", and its numbers are rougher.",[169,682],{"kind":683},"personas",[33,685,686,687,690,691,694,695,698,699,702,703,706,707,710,711,714,715,718,719,722,723,726,727,730],{},"The personas changed how both models played, a lot. On Luna, tight and patient played ",[37,688],{"k":689},"persona_tight_plays"," of its hands; loose and aggressive played ",[37,692],{"k":693},"persona_aggressive_plays",". The talker spoke on ",[37,696],{"k":697},"persona_talker_talks"," of its moves, against ",[37,700],{"k":701},"persona_control_talks"," for plain Luna. The reader cited earlier hands in ",[37,704],{"k":705},"persona_reader_history"," of its reasoning, against ",[37,708],{"k":709},"persona_control_history"," for plain Luna. Sol moved just as far: its seats played anywhere from ",[37,712],{"k":713},"persona_sol_plays_min"," to ",[37,716],{"k":717},"persona_sol_plays_max"," of their hands, its talker spoke on ",[37,720],{"k":721},"persona_sol_talker_talks"," of its moves against ",[37,724],{"k":725},"persona_sol_control_talks"," for plain Sol, and its reader cited earlier hands in ",[37,728],{"k":729},"persona_sol_reader_history"," of its reasoning.",[33,732,733,734,737,738,741,742,745,746,504,749,752],{},"They didn't measurably change the results on either model. Plain Luna took ",[37,735],{"k":736},"persona_control"," against a fair share of ",[37,739],{"k":740},"persona_fair",", and no persona beat it: the closest finished ",[37,743],{"k":744},"persona_best_behind"," behind, well inside what luck explains. Loose and aggressive won the most tournaments (",[37,747],{"k":748},"persona_aggressive_wins",[37,750],{"k":751},"persona_control_wins"," for plain Luna) with about the same average prize: more ups and downs, with no gain on average.",[33,754,755,756,759,760,763,764,767,768,771,772,774,775,714,778,781],{},"Plain Sol took ",[37,757],{"k":758},"persona_sol_control",", and ",[37,761],{"k":762},"persona_sol_below"," of the five personas finished below it. Loose and aggressive came out ",[37,765],{"k":766},"persona_sol_best_lead"," ahead, but luck alone could put that gap anywhere from ",[37,769],{"k":770},"persona_sol_best_range"," (a minus means plain Sol ahead). With only ",[37,773],{"k":676}," tournaments the ranges are wide: plain Sol's own share could sit anywhere from ",[37,776],{"k":777},"persona_sol_control_lo",[37,779],{"k":780},"persona_sol_control_hi",". A test this small can only catch a large effect, so a modest one on Sol would slip past it. What it does show is the same pattern as Luna. On both models, a persona changed how the AI played and left how well it did about where it was.",[80,783,785],{"id":784},"what-this-says-about-ai","What this says about AI",[33,787,788,791],{},[88,789,790],{},"Two very different models finished level."," Sol and Opus play differently (Opus plays far more hands and talks far more) and ended with the same prize share. A ranking that puts one of them first would be reading noise.",[33,793,794,797,798,800,801,803],{},[88,795,796],{},"The most expensive model finished last of the three guests."," Fable costs about ",[37,799],{"k":47}," times as much per move as Opus and won none of its tournaments on the same cards where Opus won ",[37,802],{"k":346},". Price per move told me nothing about who would win.",[33,805,806,809],{},[88,807,808],{},"In real games, words barely registered."," Fable and Haiku built reads on opponents from memory, such as Fable deciding Jev calls almost every bet. Almost no model's reasoning weighed what an opponent said. When I tested talk on purpose, the two players it moved took it at face value.",[33,811,812,815],{},[88,813,814],{},"A persona changed the style and left the results alone, on a cheap model and on the strongest one."," Written instructions moved how often Luna and Sol played, talked and remembered, without moving their prize share. Sol's test was small, so it confirms the direction more than it pins down the size.",[33,817,818,821,822,824],{},[88,819,820],{},"What a model can see is part of the test."," These AIs knew the payouts, saw who had been knocked out and remembered recent hands, and Opus brings up the prize money in ",[37,823],{"k":466}," of its reasoning. A test that hid those things would be measuring a different game. Before trusting any ranking of AI models, at poker or anything else, check what each one was shown.",[80,826,828],{"id":827},"sit-in-yourself","Sit in yourself",[33,830,831],{},"Take a seat against four of the five live-table AIs, and pick who sits out.",[263,833],{"mode":834},"play",[80,836,838],{"id":837},"methods","Methods",[840,841,842,872,891,897,906,924,934,943,976,982],"ul",{},[843,844,845,848,849,851,852,854,855,857,858,860,861,864,865,714,868,871],"li",{},[88,846,847],{},"Main study:"," three tables of five, one per guest (Sol, Opus, Fable), each against Luna, Haiku, Luna Decisions and Jev. All three tables used the same ",[37,850],{"k":282}," sets of decks, each played five times with the seats rotated, so the guests held identical cards. ",[37,853],{"k":177}," tournaments, ",[37,856],{"k":77}," hands, ",[37,859],{"k":293}," decisions, ",[37,862],{"k":863},"fallbacks"," timeouts or failed answers. Every tournament played to a finish (",[37,866],{"k":867},"shortest",[37,869],{"k":870},"longest"," hands).",[843,873,874,877,878,880,881,884,885,887,888,890],{},[88,875,876],{},"Format:"," ",[37,879],{"k":181}," chips each (",[37,882],{"k":883},"start_bb"," big blinds), blinds doubling every ",[37,886],{"k":189}," hands, prizes ",[37,889],{"k":197}," for the top three. Every language model used the same low reasoning setting; decision models have no such setting.",[843,892,893,896],{},[88,894,895],{},"Information:"," each AI saw its own cards, the shared cards, every player's chips and bets, how the pot splits when someone is all-in for less, who acts after it, the payouts and who has been knocked out, plus its memory of the last one to two rounds of hands. No odds, hand names or stats.",[843,898,899,902,903,905],{},[88,900,901],{},"Scoring:"," prize share, the payout each player was told it was playing for. The ranges are 95% intervals over tournaments. With ",[37,904],{"k":286}," tournaments per guest they're wide and rough: they say which results are clear (Jev last), which just clear luck (Opus over Fable) and which don't (Sol versus Opus, the leaders versus Luna).",[843,907,908,911,912,290,914,917,918,290,920,923],{},[88,909,910],{},"Personas:"," six seats, five personas and one control, on the same rotated decks twice: GPT-6 Luna for ",[37,913],{"k":663},[37,915],{"k":916},"persona_hands"," hands, and GPT-6 Sol for ",[37,919],{"k":676},[37,921],{"k":922},"persona_sol_hands"," hands.",[843,925,926,929,930,933],{},[88,927,928],{},"Data check:"," I went back through every recorded move for answers cut off by the length limit, failed answers and timeouts. The count across all three was ",[37,931],{"k":932},"fallbacks_all",", so no result needed a rerun.",[843,935,936,877,939,942],{},[88,937,938],{},"Talk test:",[37,940],{"k":941},"talk_calls"," replayed decisions from each AI's own recorded spots, with the bet set to three quarters of the pot.",[843,944,945,877,948,951,952,955,956,959,960,963,964,967,968,971,972,975],{},[88,946,947],{},"Cost:",[37,949],{"k":950},"cost_a"," for the main study (Fable's table alone ",[37,953],{"k":954},"cost_fable_table",", Opus's ",[37,957],{"k":958},"cost_opus_table",", Sol's ",[37,961],{"k":962},"cost_sol_table","), ",[37,965],{"k":966},"cost_b"," for the Luna persona table, ",[37,969],{"k":970},"cost_b_sol"," for the Sol persona table and ",[37,973],{"k":974},"cost_talk"," for the talk test, on one paid API key.",[843,977,978,981],{},[88,979,980],{},"Replays:"," the recorded hands above come straight from the run logs. Cards, talk and reasoning are as the models wrote them, trimmed to a sentence.",[843,983,984,987,988,995],{},[88,985,986],{},"Code:"," the game, the tournament runner and the analysis are open source at ",[989,990,994],"a",{"href":991,"rel":992},"https:\u002F\u002Fgithub.com\u002FGKjohns\u002Fholdem-ai",[993],"nofollow","github.com\u002FGKjohns\u002Fholdem-ai",", including the full results write-up.",{"title":997,"searchDepth":998,"depth":998,"links":999},"",2,[1000,1001,1002,1003,1004,1012,1013,1014,1015,1016],{"id":82,"depth":998,"text":83},{"id":204,"depth":998,"text":205},{"id":268,"depth":998,"text":269},{"id":312,"depth":998,"text":313},{"id":440,"depth":998,"text":441,"children":1005},[1006,1008,1009,1010,1011],{"id":451,"depth":1007,"text":452},3,{"id":476,"depth":1007,"text":477},{"id":514,"depth":1007,"text":515},{"id":544,"depth":1007,"text":545},{"id":570,"depth":1007,"text":571},{"id":593,"depth":998,"text":594},{"id":653,"depth":998,"text":654},{"id":784,"depth":998,"text":785},{"id":827,"depth":998,"text":828},{"id":837,"depth":998,"text":838},"2026-10-10","Seven AIs played Texas Hold'em on identical cards. The two leaders couldn't be told apart, the priciest model never won, and in their games almost none of them weighed a word anyone said. Step through their hands, watch them play, or take a seat.","md","cards","\u002Fblog\u002Fai-poker-night\u002Fog.png",{},"\u002Fblog\u002Fai-poker-night",{"title":1025,"description":1026,"image":1021},"AI poker night: what happened when seven AIs played Hold'em | Kyle Johnson","Five language models and two decision models played poker tournaments on identical cards. GPT-6 Sol and Claude Opus 5.5 finished level, Claude Fable 5.1 won none of its ten, and table talk barely registered. Watch them play, or take a seat.","blog\u002Fai-poker-night","_cKc1ZttzHe_mLksAZ5OzgiKTW2Pj5HMlMtMl3X9jPg",1791677345420]