Risk and Portfolio OptimizationAdvanced Bandit Algorithms ResearchConstraint Satisfaction and Optimization

Qing Zhu, Baoqin Xie, Yuze Li, Jiamiao Zhao

2026.1.1Data Science and Management

DOI: 10.1016/j.dsm.2026.01.003

Abstract

In recent years, rising geopolitical tensions, trade protectionism, and demand for safe-haven assets have driven a notable surge in financial market volatility, exposing the limitations of both traditional static and prediction-driven investment strategies. Existing reinforcement learning (RL) studies typically adopt fixed-reward designs, which limit their adaptability in non-stationary markets. To address this gap, this paper proposes a RL framework, VMD-PPO, that integrates variational mode decomposition (VMD) with proximal policy optimization (PPO). The framework leverages VMD to extract multi-scale market signals, employs PPO to optimize trading policies, and introduces a market-state-adaptive reward mechanism that dynamically balances return, risk, and transaction costs, overcoming the rigidity of static rewards in prior research. An empirical study centered on portfolio construction demonstrates that the VMD-PPO framework achieves superior performance across key metrics, including returns, maximum drawdown ( MDD ), and risk-adjusted returns, while consistently exhibiting strong robustness under various market trends. Unlike prior research, which primarily focuses on prediction accuracy, this study emphasizes adaptability and deployability, providing both a foundation and a practical pathway for applying RL to highly volatile financial assets.

Citation format

ZHU, Qing, et al. Breathing rewards: Dynamic market-aware portfolio optimization. Data Science and Management, 2026.