基于物理感知注意力MADDPG算法的集群协同控制(2026科协年会论文)

  • 郭宇航 ,
  • 王俱博玺 ,
  • 布树辉 ,
  • 韩鹏程 ,
  • 陈子建
展开
  • 西北工业大学

收稿日期: 2026-03-12

  修回日期: 2026-07-13

  网络出版日期: 2026-07-20

基金资助

国家自然科学基金;国家自然科学基金;国家资助博士后研究人员计划

Cooperative Swarm Control Based on Physics-Aware Attention MADDPG

  • GUO Yu-Hang ,
  • WANG Ju-BoXi ,
  • BU Shu-Hui ,
  • HAN Peng-Cheng ,
  • CHEN Zi-Jian
Expand

Received date: 2026-03-12

  Revised date: 2026-07-13

  Online published: 2026-07-20

摘要

针对复杂未知威胁环境下固定翼无人机集群的编队保持与协同避险问题,提出了一种基于物理感知注意力的多智能体深度确定性策略梯度(PA-MADDPG)算法。首先,设计了一个物理感知注意力模块,将人工势场法作为物理先验知识与注意力网络相结合,显著提升了集群在复杂动态威胁下的避险鲁棒性与跨规模拓扑泛化能力;其次,设计了一种自适应门控复合稠密奖励函数,有效缓解了长时域任务中稀疏奖励导致的梯度信号不足和多目标冲突问题。此外,提出了一种进化策略同步机制,通过周期性筛选高适应度智能体并将其策略参数同步到其他智能体上,结合优先经验回放机制,加速了训练收敛。仿真结果表明,在静态、动态威胁场景下,采用该算法的集群机间避障率与环境避险率均达到92%以上,任务综合效能较其他算法有明显提升;在动态拓扑泛化场景下,算法的避障与避险率仍保持在82%以上,证实了该算法不仅实现了固定翼无人机集群的高效避险与编队保持,更具备在复杂变阵任务中灵活迁移的强鲁棒性。

本文引用格式

郭宇航 , 王俱博玺 , 布树辉 , 韩鹏程 , 陈子建 . 基于物理感知注意力MADDPG算法的集群协同控制(2026科协年会论文)[J]. 航空学报, 0 : 1 -0 . DOI: 10.7527/S1000-6893.2026.33568

Abstract

To address the problems of formation maintenance and cooperative obstacle avoidance for fixed-wing unmanned aerial vehicle (UAV) swarms in complex and unknown threat environments, a physics-aware attention multi-agent deep deterministic policy gradient (PA-MADDPG) algorithm is proposed in this paper. First, a physics-aware attention module is designed, which explicitly integrates the artificial potential field (APF) method into the attention network as physical prior knowledge. This significantly enhances the collision avoidance robustness and cross-scale topological generalization capability of the swarm under complex dynamic threats. Second, an adaptive gated composite dense reward function is formulated to effectively alleviate the insufficient gradient signals caused by sparse rewards and multi-objective conflicts in long-horizon tasks. Furthermore, an evolutionary strategy synchronization mechanism is introduced. By periodically screening high-fitness agents and synchronizing their policy parameters to other agents via soft updates, coupled with prioritized experience replay (PER), the training convergence is substantially accelerated. Simulation results demonstrate that in both static and dynamic threat scenarios, the proposed algorithm achieves inter-UAV collision avoidance and environment threat avoidance rates of over 92%, significantly outperforming baseline algorithms in comprehensive task effectiveness. In dynamic topological generalization scenarios, the avoidance rates stably remain above 82%. These results confirm that the PA-MADDPG algorithm not only realizes efficient obstacle avoidance and formation maintenance for fixed-wing UAV swarms but also exhibits strong robustness for flexible transfer in complex reconfiguration tasks.
Options
文章导航

/