Sangkeum Lee, Beomdo Park, Junseong Park, Hyeonseok Jang, Hoon Jeong, Taewook Heo
Abstract
Quantum batteries promise ultrafast energy storage but are highly sensitive to noise, drift, and hardware constraints, making safe high-performance charging a central challenge for noisy intermediate-scale quantum (NISQ) devices. We propose a <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">measurement-informed safe control</i> framework that couples <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">harmonic-spectrum-based syndrome diagnostics</b>—<inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$H_{2}/H_{1}$</tex-math></inline-formula>, <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$H_{3}/H_{1}$</tex-math></inline-formula>, and frequency drift—with a <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">battery management system (BMS)-constrained curriculum reinforcement learning (RL)</b> policy. Spectral features are compressed into a three-level syndrome code (<inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$s\in \lbrace 0,1,2\rbrace$</tex-math></inline-formula>) that serves as a real-time hardware risk proxy for the controller. Our digital-twin simulator incorporates <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$T_{1}/T_\phi$</tex-math></inline-formula> relaxation, crosstalk, collective effects, and terminal-voltage dynamics, while safety risks are explicitly encoded as BMS-related penalties (state-of-health, voltage limits, and high-risk operation ratio) in the RL reward. Across staged curricula of increasing system complexity, the learned policy empirically traces a strictly improved Pareto frontier between final ergotropy and high-risk ratio compared to baseline and threshold-grid control strategies, with gains confirmed by multi-seed statistical confidence intervals. To support near-term deployment, we position the current work as a digital-twin stage and outline a concrete simulation-to-real protocol: fix ROC-calibrated thresholds, re-tune <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$(\tau ^{w},\tau ^{h})$</tex-math></inline-formula> on a small hardware calibration split, and validate a one-step voltage shield. We further demonstrate the framework on a benchtop transmon setup with <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$N=1$</tex-math></inline-formula>–2, reporting shield trigger/violation rates, sim-to-real drift of spectral features (KL/EMD), and an end-to-end latency within 20<inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$\mu$</tex-math></inline-formula>s, indicating that harmonic-syndrome-informed safe RL is a viable route toward practical quantum battery charging control.
Citation format
LEE, Sangkeum, et al. Measurement-informed safe reinforcement learning for quantum battery charging via harmonic-syndrome diagnostics and BMS constraints. IEEE Transactions on Quantum Engineering, 2026, 7: 1–15.