The merge puzzle community has long treated 2048 as the gold standard for afterstate learning, but this work reveals how quickly that playbook frays when the action space widens and chance events get a preview window. Developers are not chasing a gimmick; they are wrestling with a structural problem that will feel familiar to anyone who has tried to scale reinforcement learning beyond toy grids. The core tension is right there in the numbers: 30 legal actions, 128 simulations per real action, and a chance node that branches over six-tile previews. That budget disappears fast. The data shows it, forcing every root action to receive two visits already burns nearly half the tree, and bumping simulations from 128 to 192 changed nothing. This is not a tuning issue; it is a sign that the search architecture itself is fighting the problem's shape.
The decision to separate deterministic afterstates from explicit chance nodes is the right instinct, and it aligns with the [Planning/RL for a stochastic single-player merge puzzle: afterstates, previewed chance events, and long-horizon throughput [D]](/post/planning-rl-for-a-stochastic-single-player-merge-puzzle--afterstates--previewed-chance-events--and-long-horizon-throughput--d--cmueex7bt02vf4yswcr5rrgsi) approach that treats the preview as a conditional decision point rather than a raw stochastic event. But experiments reveal a deeper issue: the learned value head, trained on long-horizon returns, becomes unreliable when asked to evaluate leaf states after just a few actions. That is why exhaustive maximization over preview-conditioned fourth actions blew up, and why policy-only distillation saturated. The lesson is not that neural value functions are useless here; it is that they need to be paired with a search that can correct for their biases, not amplify them. The comparison to the Cloudflare's Data Innovation Frees 100 TB, Boosts DNS Performance redesign is apt in spirit, both are about reducing the cost of a bottleneck, whether that bottleneck is memory footprint or simulation depth.
The most practical takeaway for anyone building similar agents is the shift toward average-reward thinking. The 30-minute objective is not a discounted episodic problem with a fixed horizon; it is a throughput problem with a restart option. That changes what "good" means. A policy that scores 13 9s in 272 actions looks great on a per-episode basis, but it is useless if it dies quickly after the first death. The tracking of first-9 cost, subsequent-9 gaps, and survival length is exactly the right vocabulary. The gap between 35.2 9s per 1,000 actions and the human-observed 115 per 1,800 actions is not a small margin to close; it is a signal that the current value head and search are not capturing the long-horizon structure of mature boards. The idea to fit an HMM to the drop distribution is not a side quest, it is a direct response to the human-reported alternation between simple and mixed drops, and it could convert a noisy chance node into a predictable regime that the planner can exploit.
The honest answer to the central question is that no established algorithm will solve this out of the box, because the combination of afterstates, previews, and stack constraints sits at an awkward intersection of 2048-style TD, SameGame-style search, and inventory planning. The most promising thread is the suggestion of Gumbel sequential halving at the root, because it directly addresses the allocation problem without forcing a handcrafted heuristic. The second is tree reuse across observed chance outcomes, which could turn the preview from a one-shot decision point into a persistent branch that carries information across cycles. If the author pairs that with a distributional value head that predicts 9-count quantiles over multiple horizons, they may finally get a signal that does not collapse under its own variance. The specific detail to watch is the death-risk penalty: at 0.5, it is already shaping behavior, but the author has not yet shown how it interacts with the average-reward objective. That interaction, more than any single algorithmic trick, will determine whether the 115-9s target is reachable or just a number on a leaderboard.