Sangkeum Lee, Beomdo Park, Junseong Park, Hyeonseok Jang, Hoon Jeong, Taewook Heo

2026IEEE Transactions on Quantum Engineering

DOI: 10.1109/tqe.2026.3670136

Abstract

Quantum batteries promise ultrafast energy storage but are highly sensitive to noise, drift, and hardware constraints, making safe high-performance charging a central challenge for noisy intermediate-scale quantum (NISQ) devices. We propose a <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">measurement-informed safe control</i> framework that couples <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">harmonic-spectrum-based syndrome diagnostics</b>—<inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$H_{2}/H_{1}$</tex-math></inline-formula>, <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$H_{3}/H_{1}$</tex-math></inline-formula>, and frequency drift—with a <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">battery management system (BMS)-constrained curriculum reinforcement learning (RL)</b> policy. Spectral features are compressed into a three-level syndrome code (<inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$s\in \lbrace 0,1,2\rbrace$</tex-math></inline-formula>) that serves as a real-time hardware risk proxy for the controller. Our digital-twin simulator incorporates <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$T_{1}/T_\phi$</tex-math></inline-formula> relaxation, crosstalk, collective effects, and terminal-voltage dynamics, while safety risks are explicitly encoded as BMS-related penalties (state-of-health, voltage limits, and high-risk operation ratio) in the RL reward. Across staged curricula of increasing system complexity, the learned policy empirically traces a strictly improved Pareto frontier between final ergotropy and high-risk ratio compared to baseline and threshold-grid control strategies, with gains confirmed by multi-seed statistical confidence intervals. To support near-term deployment, we position the current work as a digital-twin stage and outline a concrete simulation-to-real protocol: fix ROC-calibrated thresholds, re-tune <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$(\tau ^{w},\tau ^{h})$</tex-math></inline-formula> on a small hardware calibration split, and validate a one-step voltage shield. We further demonstrate the framework on a benchtop transmon setup with <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$N=1$</tex-math></inline-formula>–2, reporting shield trigger/violation rates, sim-to-real drift of spectral features (KL/EMD), and an end-to-end latency within 20<inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$\mu$</tex-math></inline-formula>s, indicating that harmonic-syndrome-informed safe RL is a viable route toward practical quantum battery charging control.

Citation format

LEE, Sangkeum, et al. Measurement-informed safe reinforcement learning for quantum battery charging via harmonic-syndrome diagnostics and BMS constraints. IEEE Transactions on Quantum Engineering, 2026, 7: 1–15.