No-limit hold'em · decision models vs. perfect play and vs. bots
100% vibe-coded · orchestrated by Claude Opus 5.5
Can decision models play poker?
Loading results…
Insights · the key findings
What we learned
Fifteen things we found by making AI models play poker against a perfect-play solver and against bots. Each card links to the evidence.
Each card shows the exact statistic, sample size and 95% CI where available. Cards marked live are computed from the current export (bench.json); the others come from data/insights.json. Every card is linkable (#insight-…).
The short version
What we found
Featured finding · confidence
The models know when they're unsure
Every answer comes with a probability for each move. We grouped the answers by how much the model put on its favourite move — how sure it was — and checked how often that move was really the best one, and what the mistakes cost.
How sure the model was → how often it picked the best move
Reliability diagram
Stated confidence vs. how often it was right
x = mean top probability in the bin, y = share of picks within 0.25% pot of the best EV. Dashed diagonal = perfectly calibrated. Bins with fewer than 20 answers are hidden.
Escalation: a second opinion only when unsure
Every model, every bucket
By prompt variant
Part 1 · solved spots
The spot test
Each model saw the same turn and river situations and picked a move. We compared every pick with the move a poker solver (a perfect-play computer) would make.
Points kept, out of 100
100 = always picks the best move. The bar around each dot shows how sure we are (95% range). Gold lines are simple bots, for scale.
Leaderboard
Ranked by GTO score. Every decision model is scored twice: on the action it picked (argmax) and on the full probability distribution it returned (mixed). Reference bots are shown in gold for scale. Click a column heading to sort.
Six views of the same table
GTO score, with 95% confidence intervals
Dot = mean score; whisker = 95% CI. Vertical lines are the reference bots.
Picking vs. playing the distribution
EV lost per decision, % of pot. Argmax = the model's pick; mixed = sampling its probabilities.
Price vs. play
GTO score against cost per 1,000 decisions (log scale). Up and left is better.
Turn vs. river
GTO score by street.
Leading vs. facing a bet
GTO score when checking/betting first vs. when facing a bet.
Latency per decision
Median response time (dot) and 90th percentile (end of line), seconds, log scale.
Part 1b · five models, one decision
When all five agree
We asked five different AI models the same questions and looked at how many of them chose the same move — and whether the move most of them chose was any good.
How many models agreed → what the shared pick cost
Majority vote vs. the best single model
Average share of the pot given away per decision. Shorter is better.
Per model on the same decisions
Bet sizing: share of each model's loss from the right action at the wrong size
Every suite
Part 2 · full hands, every format
Real hands, cash to tournaments
Here the models play whole hands — every decision from the flop to the river, or every decision before the flop — in the formats people actually play. We count how many big blinds they give away per 100 hands compared with perfect play. Lower is better; 0 is perfect.
Postflop suites: flops solved from each format's preflop ranges, hands dealt and played to the end along the GTO line, every decision of both seats recorded. Preflop suites: heads-up blind-vs-blind / Spin HU trees from the preflop solver. Metric: bb/100 lost vs GTO = 100 × mean over hands of Σ EV loss ÷ 2 (the model plays both seats), 95% CI over hands. Hands shared by a suite and its dev/test splits are counted once.
Every player, every format
Suites
Part 3 · the arena
Heads-up against real opponents
Perfect play is one yardstick; real opponents are another. Each model played hundreds of complete heads-up hands against two bots: a solid regular that plays tight and aggressive, and a calling station that never folds. Every deal is played twice with the seats swapped, so luck of the cards mostly cancels out. Higher is better; above 0 means the model won.
Duplicate heads-up matches, 100bb, every deal played twice with seats swapped. bb/100 ± 95% CI over duplicate pairs (pooled across matches of the same pairing); all-in EV bb/100 removes all-in luck. Matches with fewer than 100 hands and self-play are excluded; bot-vs-bot matches are references.
Decision models vs. bots
Bot vs. bot references
Part 4 · experiments
Pushing GPT-6 Luna
Can a model get better just by asking differently? We tried several ways of describing the table and asking the question. Each idea was tuned on one set of practice hands, then checked on fresh hands it had never seen — the only result that counts.
Variants of the request (state and question) and post-hoc players derived from stored answers, on the dev split (tuning) and the held-out test split (confirmation) of the cash hands suite. Δ = paired difference in bb/100 lost vs the default variant (v2) over the hands both answered, 95% CI. Negative is better. Hybrid players use non-model heuristics and are not pure model results.
Dev and test, side by side
Every hand, every answer
Spot explorer
Pick a suite and a decision to see the table as the models saw it, the solver's strategy for the hero's exact hand, and what each model would have done. Links like #d042 (spot suite) or #cash-hands-dev/d042 open a decision directly. Use J / K to step through the list.
How the numbers are made