Bo Xia, Yilun Kong, Yongzhe Chang, Bo Yuan, Zhiheng Li, Xueqian Wang, Bin Liang

2026IEEE Transactions on Artificial Intelligence

DOI: 10.1109/tai.2026.3664781

Abstract

Real-world decision-making systems often suffer from non-negligible observation and action delays arising from sensing and actuation, which break the standard Markov assumption and degrade reinforcement learning (RL) performance. We present <underline xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">D</u>elay-resilient <underline xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">E</u>ncoder-<underline xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">E</u>nhanced <underline xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">R</u>L (DEER), a practical framework that systematically integrates well-established techniques (sequence-to-sequence representation learning, information-state construction, and offline-to-online RL) to enable robust performance in delayed environments. DEER first pretrains an encoder on offline delay-free trajectories and then uses it online to map delayed observations with recent action histories into fixed-length context vectors for standard RL agents. We further provide a theoretical decomposition that links reconstruction error, distributional mismatch, latent-policy generalization, and environment mismatch to overall performance, offering interpretable guidance for design choices. Empirically, integrating DEER with Soft Actor-Critic yields competitive and stable performance across a range of constant and random delaysettings on six continuous-control benchmarks. The contribution is therefore twofold: (i) a practical, deployable engineering solution for delayed RL that leverages existing building blocks effectively; and (ii) a quantitative analysis that explains when and why this integration works.

Citation format

XIA, Bo, et al. Delay-aware reinforcement learning with encoder-enhanced state representations. IEEE Transactions on Artificial Intelligence, 2026.