Xing Fang, Xiaoli Luan, Feng Ding, Fei Liu
2026.1.8ASIAN JOURNAL OF CONTROL
tlooto Summary
Two examples involving a flight control system and a continuous fermentation reactor system are simulated, demonstrating that the proposed methods can effectively enhance the speed of the Q function estimation in the initial stage and achieve the optimal control for the target task within less time steps.
Abstract
In policy iteration based on Q‐learning, both a stable initial control policy and sufficient time steps are essential to estimate the Q function precisely, which accordingly affects the optimal control performance. To improve the estimate speed in the initial stage of policy iteration, an information‐matrix‐transfer strategy is embedded in the policy iteration based on the Q‐learning. Here, we reuse the information matrix of a source task, which serves as a foundation for selecting the initial control policy and estimating the target Q function through constructing an information‐matrix‐transfer estimation equation. Meanwhile, the convergence of the proposed method is analyzed. Moreover, a probabilistic information‐matrix‐transfer strategy is introduced to reduce negative transfer occurrences. Finally, two examples involving a flight control system and a continuous fermentation reactor system are simulated, demonstrating that the proposed methods can effectively enhance the speed of the Q function estimation in the initial stage and achieve the optimal control for the target task within less time steps.
Citation format
FANG, Xing, et al. Policy iteration for model‐free optimal control based on both q‐learning and information‐matrix‐transfer. ASIAN JOURNAL OF CONTROL, 2026.