Computer Science

Weikai Wang, Kiminori Matsuzaki

2026.3.1IEEE Transactions on Games

DOI: 10.1109/tg.2025.3606517

Abstract

The game <italic>2048</italic> has attracted millions of people with its simple yet challenging gameplay, leading to the development of numerous computer players. Most successful computer players for <italic>2048</italic> use evaluation functions trained through temporal difference learning (TD learning) or its variants. While TD learning is highly effective and can improve evaluation functions quickly, the performance of these functions often plateaus after a certain number of timesteps. Therefore, it is important to refine those evaluation functions to further enhance the performance of computer players. In this article, we extend the conventional TD learning approach and propose two refinement algorithms for <italic>2048</italic>. First, we conducted detailed experiments to refine the best open-source neural network, and achieved significant performance improvements, increasing the average score from <inline-formula><tex-math notation="LaTeX">$2.49 \times 10^{5}$</tex-math></inline-formula> to <inline-formula><tex-math notation="LaTeX">$3.37 \times 10^{5}$</tex-math></inline-formula> in greedy play (1-ply lookahead) and from <inline-formula><tex-math notation="LaTeX">$4.87 \times 10^{5}$</tex-math></inline-formula> to <inline-formula><tex-math notation="LaTeX">$5.45 \times 10^{5}$</tex-math></inline-formula> with 3-ply expectimax search. We also applied our refinement method to the state-of-the-art N-tuple network, improving the average score from <inline-formula><tex-math notation="LaTeX">$5.85\times 10^{5}$</tex-math></inline-formula> to <inline-formula><tex-math notation="LaTeX">$6.10\times 10^{5}$</tex-math></inline-formula> with 6-ply expectimax search and the tile-downgrading trick.

Citation format

WANG, Weikai; MATSUZAKI, Kiminori. Refining evaluation functions for game 2048 by extended temporal difference learning. IEEE Transactions on Games, 2026, 18(1): 56–65.