首页 >

部分可观环境下多无人机通信覆盖路径规划方法

左燕1,林韦屹2,彭冬亮1,黄猛2   

  1. 1. 杭州电子科技大学
    2. 杭州电子科技大学自动化学院
  • 收稿日期:2026-05-20 修回日期:2026-09-09 出版日期:2026-09-20 发布日期:2026-09-20
  • 通讯作者: 林韦屹
  • 基金资助:
    国家自然科学基金;国家科技重大专项

Multi-UAV path planning for communication coverage in partially observable scenario

  • Received:2026-05-20 Revised:2026-09-09 Online:2026-09-20 Published:2026-09-20
  • Contact: Wei-Yi LIN

摘要: 随着无人机技术的发展,无人机通信基站逐渐成为灾后应急通信快速响应的重要保障。在灾后应急通信场景中,多无人机需要在有限的电池容量和环境部分可观条件下,通过分布式协同无人机充电和飞行决策实现受灾区域内运动用户的通信覆盖率最大化、覆盖公平性以及能量消耗最小化联合优化。对此,本文建立基于去中心化部分可观测马尔可夫决策的多机协同调度模型,提出了一种基于预测辅助的双深度循环图网络(Predictive Double Deep Graph Network,PDDGN)的算法。该算法通过构建预测网络学习环境动态转移规律来估计未来状态,以补偿部分可观条件下局部观测的信息缺失并适应场景的动态性。同时引入双目标网络结构,在训练阶段通过双目标网络差异化软更新以提高算法训练的稳定性和收敛效率。仿真结果表明,所提PDDGN算法在部分可观环境下的动态通信覆盖任务中能有效提高无人机机群通信覆盖率和公平性、降低能耗,和传统深度强化学习算法相比综合覆盖指标覆盖-公平性-能量(Coverage-Fairness-Energy,CFE)得分最大提高7%,收敛回合数最大缩短60.9%,并具有良好的稳定性与可拓展性。

关键词: 路径规划, 多无人机, 应急通信, 多目标优化, 深度强化学习

Abstract: With the advancement of unmanned aerial vehicle (UAV) technology, UAV-enabled aerial communication base stations have increasingly become a crucial means of rapidly restoring emergency communications in post-disaster scenarios. In post-disaster emergency communication scenarios, multiple UAVs need to jointly optimize communication coverage maximization, coverage fairness, and energy consumption minimization for moving users in the affected area under limited battery capacity and partial observability, through distributed collaborative decision-making on UAV charging and flight. To address this, we formulate a cooperative multi-agent scheduling model based on a decentralized partially observable Markov decision process (Dec-POMDP). We then propose a predictive-assisted algorithm termed the Predictive Double Deep Graph Network (PDDGN). Specifically, PDDGN incorporates a prediction network to estimate future state observations and employs a deep recurrent graph network to integrate current observations with predicted information, thereby learning the observation–action value functions for UAVs and adapting to environmental dynamics. In addition, a double-target-network architecture is introduced; during training, differentiated soft updates between the two target networks are used to enhance training stability and improve learning efficiency. Simulation results demonstrate that, for dynamic communication coverage tasks under partial observability, the proposed PDDGN effectively improves communication coverage and fairness while reducing energy consumption. Compared with conventional deep reinforcement learning baselines, PDDGN achieves up to a 7% improvement in the composite Coverage–Fairness–Energy (CFE) score and a maximum reduction of 60.9% in convergence episodes, while exhibiting good stability and scalability.

Key words: Path planning, Multi-Unmanned Aerial Vehicle(UAV), Emergency communications, Multi-objective optimization, Deep Reinforcement Learning(DRL)

中图分类号: