Equilibriumbench

No-limit hold'em · decision models vs. perfect play and vs. bots

100% vibe-coded · orchestrated by Claude Opus 5.5

Can decision models play poker?

Loading results…

Insights · the key findings

What we learned

Fifteen things we found by making AI models play poker against a perfect-play solver and against bots. Each card links to the evidence.

Each card shows the exact statistic, sample size and 95% CI where available. Cards marked live are computed from the current export (bench.json); the others come from data/insights.json. Every card is linkable (#insight-…).

    Part 1 · solved spots

    The spot test

    Each model saw the same turn and river situations and picked a move. We compared every pick with the move a poker solver (a perfect-play computer) would make.

    Points kept, out of 100

    100 = always picks the best move. The bar around each dot shows how sure we are (95% range). Gold lines are simple bots, for scale.

    Leaderboard

    Ranked by GTO score. Every decision model is scored twice: on the action it picked (argmax) and on the full probability distribution it returned (mixed). Reference bots are shown in gold for scale. Click a column heading to sort.

    Spot benchmark leaderboard. Columns are sortable.

    Six views of the same table

    GTO score, with 95% confidence intervals

    Dot = mean score; whisker = 95% CI. Vertical lines are the reference bots.

    Picking vs. playing the distribution

    EV lost per decision, % of pot. Argmax = the model's pick; mixed = sampling its probabilities.

    Price vs. play

    GTO score against cost per 1,000 decisions (log scale). Up and left is better.

    Turn vs. river

    GTO score by street.

    Leading vs. facing a bet

    GTO score when checking/betting first vs. when facing a bet.

    Latency per decision

    Median response time (dot) and 90th percentile (end of line), seconds, log scale.

    Part 1b · five models, one decision

    When all five agree

    We asked five different AI models the same questions and looked at how many of them chose the same move — and whether the move most of them chose was any good.

    How many models agreed → what the shared pick cost

    Majority vote vs. the best single model

    Average share of the pot given away per decision. Shorter is better.

    Per model on the same decisions

    Bet sizing: share of each model's loss from the right action at the wrong size

    Every suite

    Part 2 · full hands, every format

    Real hands, cash to tournaments

    Here the models play whole hands — every decision from the flop to the river, or every decision before the flop — in the formats people actually play. We count how many big blinds they give away per 100 hands compared with perfect play. Lower is better; 0 is perfect.

    Postflop suites: flops solved from each format's preflop ranges, hands dealt and played to the end along the GTO line, every decision of both seats recorded. Preflop suites: heads-up blind-vs-blind / Spin HU trees from the preflop solver. Metric: bb/100 lost vs GTO = 100 × mean over hands of Σ EV loss ÷ 2 (the model plays both seats), 95% CI over hands. Hands shared by a suite and its dev/test splits are counted once.

    Every player, every format

    Suites

    Part 3 · the arena

    Heads-up against real opponents

    Perfect play is one yardstick; real opponents are another. Each model played hundreds of complete heads-up hands against two bots: a solid regular that plays tight and aggressive, and a calling station that never folds. Every deal is played twice with the seats swapped, so luck of the cards mostly cancels out. Higher is better; above 0 means the model won.

    Duplicate heads-up matches, 100bb, every deal played twice with seats swapped. bb/100 ± 95% CI over duplicate pairs (pooled across matches of the same pairing); all-in EV bb/100 removes all-in luck. Matches with fewer than 100 hands and self-play are excluded; bot-vs-bot matches are references.

    Decision models vs. bots

    Bot vs. bot references

    Part 4 · experiments

    Pushing GPT-6 Luna

    Can a model get better just by asking differently? We tried several ways of describing the table and asking the question. Each idea was tuned on one set of practice hands, then checked on fresh hands it had never seen — the only result that counts.

    Variants of the request (state and question) and post-hoc players derived from stored answers, on the dev split (tuning) and the held-out test split (confirmation) of the cash hands suite. Δ = paired difference in bb/100 lost vs the default variant (v2) over the hands both answered, 95% CI. Negative is better. Hybrid players use non-model heuristics and are not pure model results.

    Dev and test, side by side

    Every hand, every answer

    Spot explorer

    Pick a suite and a decision to see the table as the models saw it, the solver's strategy for the hero's exact hand, and what each model would have done. Links like #d042 (spot suite) or #cash-hands-dev/d042 open a decision directly. Use J / K to step through the list.

      How the numbers are made

      How we tested