Jin Nakazato, Hideya So, Gia Khanh Tran, Katsuya Suto
2026.6.1IEEE Journal on Miniaturization for Air and Space Systems
Abstract
In disaster scenarios, reliable emergency communication is essential for rapid medical response and public safety; however, terrestrial infrastructure is often damaged or congested. This article proposes a cooperative uncrewed aerial vehicle (UAV) ad hoc network architecture driven by multi-agent reinforcement learning (MARL) to jointly support wide-area exploration and resilient backhaul formation. We design a multiobjective reward that balances: 1) landmark discovery for expanding coverage in affected areas; 2) multihop connectivity to maintain end-to-end reachability to a ground base station (BS); and 3) mutual distance regularization to avoid excessive clustering and isolated UAVs. Using multi-agent proximal policy optimization (MAPPO), we evaluate the emergent UAV deployment behaviors under various reward-weight settings and clarify the characteristic tradeoffs between exploration and connectivity. Furthermore, after determining the UAV placements, we assess the interference-limited link quality by computing the signal-to-interference-plus-noise ratio (SINR) under multiple propagation models, including free-space, two-ray ground reflection, and 3GPP urban macro aerial vehicle (UMa-AV)/urban micro aerial vehicle (UMi-AV) LOS models. The simulation results demonstrate that appropriate reward balancing enables UAVs to reach distant targets while preserving multihop connectivity and that multipath-induced altitude sensitivity can cause severe SINR degradation in two-ray environments. These findings provide practical insights for designing learning-driven UAV-assisted emergency communication systems with robust connectivity and radio-aware deployments.
Citation format
NAKAZATO, Jin, et al. Multi-agent reinforcement learning for resilient UAV ad hoc backhaul networks. IEEE Journal on Miniaturization for Air and Space Systems, 2026, 7(2): 232–245.