Alphamon: Bit-Exact GPU Self-Play, a Training-Free Value Bound, and Search-Dominated Play
An AlphaGo-style agent for the Pokémon TCG — with bit-exact GPU self-play, a training-free ceiling on what learning can buy, and search-dominated play.
In brief
Alphamon is a complete AlphaGo-style agent for the Pokémon Trading Card Game — and three transferable lessons from building it. It ports a 1,267-card commercial engine to CUDA at one game per thread, certified bit-exact over 70,000 games and 7.9M decisions; uses that engine for a training-free bound showing the AlphaZero recipe sits near an aleatoric ceiling the game itself imposes; and finds that search, not learning, carries the agent's strength.
Key results
- Bit-exact GPU port: the full 1,267-card TCG engine (a roughly 130-opcode effect VM, turn driver, legal-action generator, featurizer) runs one game per CUDA thread and certified bit-for-bit against the reference C++ engine over 70,000 games and 7,888,623 agent decisions with zero divergences.
- Throughput: the device value stream sustains 31,951 games/s (5.05M agent-steps/s) in the batched wave loop and 89,000 games/s in a fused whole-game kernel, versus 3,208 games/s for the reference engine on one CPU core.
- A training-free ceiling: the replay-flip rate F=0.207 bounds the Bayes sign-error via 1/2 E[F] <= Err* <= E[F], bracketing the best possible value accuracy to [0.79, 0.90] rather than 1.0 -- an afternoon of engine replays forecasting a wall that the scaling program later spent weeks reaching.
- The three classical epistemic levers are jointly spent: 48x more self-play data buys +0.048 (then overfits), doubling capacity buys +0.0004, and a perfect-information oracle buys +0.003.
- The AlphaZero improvement operator is structurally inert: wiring the strongest value net into Gumbel reanalyze changes 0.00% of policy targets because 91% of reanalyzed states have every action already visited.
- Search outweighs every learned component: at byte-identical weights, replacing argmax play with PIMC-PUCT (48 sims) moves win-rate against the reference agent from 0.425 to 0.912, while no offline metric the authors could construct predicts ladder rank (rho=0.112, R^2=0.018).
The Regime Question
Alphamon is a faithful AlphaGo-style agent -- policy net, value net, Monte-Carlo search, self-play -- built for the Pokemon Trading Card Game. That recipe was proven on Go, chess, and shogi, which share two traits it silently leans on: outcomes are fully determined and fully observed, so 'who wins from state s' is a learnable function that can be driven toward {0,1}. The TCG honors neither: hands are hidden, decks are shuffled, and many effects resolve on coin flips.
The reflexive move is to classify this as a poker-family problem and reach for counterfactual regret and belief-state re-solving. The paper's central claim is that this classification is misleading. Measured on two coordinates -- a replay-flip rate E[F]=0.207 and an oracle gain Delta=+0.003 -- the game sits in the backgammon regime of future-chance variance, not the poker regime of concealed present state. That distinction decides which half of the imperfect-information literature applies and, more usefully, tells you where to spend compute before committing months to a value network.
![Two measurable coordinates -- the replay-flip rate E[F] and the oracle gain Delta -- place the TCG in the lower-right future-chance regime (backgammon's machinery) rather than the upper-left hidden-information regime where it is usually filed.](/papers/alphamon/fig_regime_map.png)
A Bit-Exact Engine On The GPU
Net-guided self-play ran at only 12-15 games/s because every node round-trips out of C++ into Python featurization and a batch-1 network call. The fix was to move the game itself onto the GPU. This is tractable because a modern card engine is not 1,267 hand-written routines but a data-oriented virtual machine: each effect is an array of Op records, the interpreter is a switch over about 130 opcodes, and state is fixed-size plain-old-data. The device state is 7,498 int32 words (about 30KB) per thread, carrying both players' hidden zones and every mid-effect transient, and it serializes bit-exactly so the device state itself is the input to the parity hash.
Correctness came before speed. A parity oracle, built before the port, hashes the complete state (via SHA-1 over a canonical serialization) at every agent decision, with unimplemented paths made loud by construction so an unported opcode cannot pass silently. The paper reports one methodological failure in full: an early featurizer gate tested for drift using a value the featurizer produced itself, so its committed artifact declared PASS after discarding all 3,000 samples as drift and certifying over none of them; it was withdrawn and rebuilt against the independently certified device legality query. The rebuilt gate certifies 70,000 games / 7,888,623 decisions with zero divergences in 35.4 minutes on one RTX PRO 4000 Blackwell.
What The Gate Caught
The certificate matters because it failed first. A triage burn-down, ranking unimplemented surface by hit rate under real play, exposed bugs no unit test would find: an ability predicate approximated as always-true (about 20 divergences per 350 games), and a single mis-identified card constant -- a tool-suppression check reading a Pokemon's ID instead of the stadium that suppresses tools -- that was the shared root cause of about 190 divergences across four unrelated symptoms. Usable device self-play rose from 87% to 94.6% to 98.84% across opcode rounds, then from 99.24% to 99.9987% (2 diverged in 150,019 games) as silent divergences were individually attributed.
The measured throughput is reported honestly by target quality, since streams are only comparable within a group. For the value stream the device wave loop hits 31,951 games/s against 3,208 for a CPU core; the fused kernel reaches 89,000 games/s. Net-guided self-play gains 4x (62->246 games/s) once the featurizer moves on-device. The residual bottleneck is now the tree itself: with engine and featurizer resident, GPU utilization is 4% and raising concurrency 4x moves throughput only 87->89 records/s, because tree search is a branching, data-dependent, per-move traversal that chases pointers node by node and does not port. The paper gives a tree-free Gumbel sequential-halving reformulation as a fixed-shape tensor program over the certified kernels.
A Training-Free Value Ceiling
The MSE-optimal value function is the conditional expectation v*(s)=2p(s)-1, and its error decomposes into a reducible epistemic term and an irreducible aleatoric term E[1-v*(s)^2] that is a property of the game. No evaluator's sign-accuracy can exceed the Bayes accuracy 1 - Err*, where Err*=E[min(p,1-p)] is zero only when outcomes are deterministic in the state -- the Go regime.
The paper's estimator makes this measurable without training anything. Re-running a fixed position under fresh randomness flips the winner with probability F(s)=2p(s)(1-p(s)), and Proposition 3 proves 1/2 E[F] <= Err* <= E[F] pointwise, so an average flip rate -- computable from engine replays alone, no network, no labels beyond outcomes -- brackets the ceiling for any value function. Measured here, E[F]=0.207 gives Err* in [0.104, 0.207] and Acc* in [0.793, 0.896]. In Go or chess the same quantity is exactly zero. This is the single measurement that should have preceded the entire value-scaling program.

Every Lever Is Spent
Against that [0.79, 0.90] ceiling, the three classical epistemic reducers were driven to exhaustion with same-data controls. Data: an 855k-game corpus (48x a 17.5k baseline) lifts held-out sign-accuracy 0.700->0.722->0.748 and then overfits -- 48x the data buys 4.8 points. Capacity: a wide network on the identical corpus scores 0.7486 against 0.7482. Information: an oracle value net trained with the opponent's hidden hand, prizes, and deck order revealed scores 0.716 against a blind net's 0.713 on the same 30k data -- perfect information is worth +0.003. When revealing every hidden card leaves accuracy unmoved, the leftover error cannot be attributed to concealment.
Crucially, the plateau at 0.748 sits below the band, so an epistemic gap remains that none of these levers moves. The resolution is that aleatoric noise both sets a floor and inflates the sample complexity of approaching it: with label variance 2E[F] approx 0.41 (measured on mid-game positions), a 48x data increase cuts statistical error only sqrt(48) approx 6.9x. Because that E[F]=0.207 is measured on a mid-game distribution rather than the value network's own evaluation distribution, the [0.79, 0.90] band is indicative rather than an exact same-distribution bound. The fix is variance-aware estimation, not a bigger corpus. And the improvement operator cannot even consume a better value: wiring the 0.748 net into Gumbel reanalyze changes 0.00% of policy targets, because the small branching factor (b-bar=7.8) means 91% of states have every action already visited and the completion term multiplies a mask that is all zeros.
![Held-out value sign-accuracy rises 0.700->0.722->0.748 as self-play data grows 48x and then overfits, saturating far below the shaded Bayes band [0.793, 0.896] -- with a wide network on the same corpus indistinguishable.](/papers/alphamon/fig_scaling.png)
Search, Not Learning
The single largest measured effect comes from search at fixed weights. Holding one behavior-cloned network byte-identical and varying only the actor against the reference rules agent, win-rate climbs from 0.425 (argmax, no search) to 0.688 (1-ply lookahead) to 0.912 (PIMC-PUCT, 48 simulations, CI [0.830, 0.957]); a weaker checkpoint replicates the ladder with cleanly separated intervals. This is the empirical shadow of the ceiling: when the value target is variance-capped, compute spent re-sampling the future dominates compute spent fitting its expectation. Consistently, the deployed leaf is lambda=1.0 pure rollout -- the head is skipped rather than faulty, because a playout returns a lower-variance Monte-Carlo estimate of v*.
Evaluation is the third result. No offline metric the authors could construct predicts the live ladder: the most rigorous harness, a clone-invariant Nash gauntlet over seven board-anchored champions across 78 pairings, attains rho=0.112 (p_perm=0.814) and R^2=0.018. The reason is that the ladder labels are not reproducible -- resubmitting an unchanged bundle settles up to 149 rating points apart (median 88 over five comparisons), exceeding the 118-point spread separating the candidates. The paper retracts one of its own conclusions (a claimed 363-point FPU refutation that was reading transient pre-convergence ratings and settled to an 89-point gap inside the noise), and reports a final standing of 4741/6807.

Abstract
Alphamon is a complete AlphaGo-style agent for the Pokémon Trading Card Game (TCG), and three results from building it that transfer beyond the game — one on systems, one on analysis, one on evaluation. Systems: we port a 1,267-card, commercial-scale TCG engine — a ~130-opcode card-effect bytecode VM with its turn driver, legal-action generator, and neural featurizer — to CUDA at one full-fidelity game per thread, and certify it bit-exact against the reference C++ engine over 70,000 games and 7,888,623 agent decisions with zero divergences, hashing the complete state at every step. The device engine sustains 31,951 games/s in a batched lockstep wave loop and 89,000 games/s in a fused whole-game kernel, against 3,208 games/s for the reference engine on one CPU core. Analysis: using the same engine we give a cheap, training-free estimator of whether the AlphaZero recipe can work on a game at all. Replaying a fixed position under fresh randomness flips the winner with probability F, and we prove ½·E[F] ≤ Err* ≤ E[F] for the Bayes sign-error Err* of any value function; measured here, F = 0.207 brackets the value-accuracy ceiling to [0.79, 0.90]. Against that ceiling the three classical epistemic levers are jointly spent: 48× more self-play data buys +0.048 and then overfits, doubling capacity buys +0.0004, and a perfect-information oracle buys +0.003. Evaluation: on a live TrueSkill-style public ladder, no offline metric we could construct predicts rank, and resubmissions of unchanged bundles settle as much as 149 rating points apart. Throughout, search — not any learned component we trained — is the dominant source of strength: holding weights fixed, replacing argmax play with PIMC-PUCT moves win-rate against the reference agent from 0.425 to 0.912.
More figures

Fig. 1The PIMC-PUCT deploy actor: a meta-archetype belief converts the partial observation into K=3 determinized worlds, PUCT (c_puct=5, prior temperature tau=2) searches each on the native engine, and root visit counts are averaged to choose the played action. 
Fig. 2AZNet: a permutation-invariant set of 36 hybrid tokens (card-ID embedding + attribute MLP + zone marker + live features) processed by a depth-3 ISAB set-transformer, with the opponent hand held in a padded zone so no hidden information can leak into the encoder. 
Fig. 3Settled ladder ratings for every net-based submission against the field's quantiles; red segments join unchanged bundles, with the champion bundle fired three times spanning 148.8 points -- label noise comparable to the entire spread of distinct agents produced.
