Qing Zhu, Baoqin Xie, Yuze Li, Jiamiao Zhao
2026.1.1Data Science and Management
Abstract
In recent years, rising geopolitical tensions, trade protectionism, and demand for safe-haven assets have driven a notable surge in financial market volatility, exposing the limitations of both traditional static and prediction-driven investment strategies. Existing reinforcement learning (RL) studies typically adopt fixed-reward designs, which limit their adaptability in non-stationary markets. To address this gap, this paper proposes a RL framework, VMD-PPO, that integrates variational mode decomposition (VMD) with proximal policy optimization (PPO). The framework leverages VMD to extract multi-scale market signals, employs PPO to optimize trading policies, and introduces a market-state-adaptive reward mechanism that dynamically balances return, risk, and transaction costs, overcoming the rigidity of static rewards in prior research. An empirical study centered on portfolio construction demonstrates that the VMD-PPO framework achieves superior performance across key metrics, including returns, maximum drawdown ( MDD ), and risk-adjusted returns, while consistently exhibiting strong robustness under various market trends. Unlike prior research, which primarily focuses on prediction accuracy, this study emphasizes adaptability and deployability, providing both a foundation and a practical pathway for applying RL to highly volatile financial assets.
Citation format
ZHU, Qing, et al. Breathing rewards: Dynamic market-aware portfolio optimization. Data Science and Management, 2026.