首页 >

基于预测增强HAPPO的空海协同拦截任务分配(飞行器协同作战技术专栏)33067重投

朱豪杰1,陈谋1,闫超2,周同乐2   

  1. 1. 南京航空航天大学自动化学院
    2. 南京航空航天大学
  • 收稿日期:2026-01-13 修回日期:2026-06-12 出版日期:2026-06-16 发布日期:2026-06-16
  • 通讯作者: 朱豪杰

Air-Sea Cooperative Missile Interception Task Allocation via Heterogene-ous-Agent PPO with Hit Probability Prediction Network

  • Received:2026-01-13 Revised:2026-06-12 Online:2026-06-16 Published:2026-06-16
  • Contact: Hao-jie ZHU

摘要: 针对协同反导拦截响应迟缓、效率低下的难题,提出一种基于异构智能体近端策略优化(HAPPO)的空海协同任务分配方法。首先,建立雷达探测跟踪、导弹动力学及毁伤效能模型,并在多智能体强化学习框架下对拦截分配问题进行形式化描述。在此基础上,为解决固定步长导致的决策冗余,设计基于事件触发的步长调整机制,实现了事件触发下的任务重构。为提升拦截任务分配策略的训练稳定性,提出了基于HAPPO的分阶段训练算法,通过渐进引入基于拦截命中率预测网络的增强奖励,提高了异构智能体的协同决策能力。仿真结果表明,所提算法有效克服了稀疏奖励的缺陷,具备空海协同任务分配能力,显著提升了对敌方反舰导弹的拦截成功率。

关键词: 任务分配, 异构智能体强化学习, 导弹拦截, HAPPO算法, 命中率预测

Abstract: To address the issues of slow response and low interception efficiency in traditional task allocation methods in co-operative missile interception, this paper proposes an air-sea cooperative missile interception task allocation meth-od based on Heterogeneous-Agent Proximal Policy Optimization (HAPPO). First, models of radar detection and tracking, missile dynamics, and damage effectiveness are established, and the missile interception allocation prob-lem is formally formulated within a heterogeneous multi-agent reinforcement learning framework. On this basis, a step adjustment mechanism based on event-triggered is developed to achieve task reconfiguration upon triggering events. To enhance the training stability of the interception task allocation policy, a staged training algorithm based on HAPPO is introduced, which progressively incorporates an enhanced reward from an interception hit-rate pre-diction network, thereby enhancing the cooperative decision-making capability of heterogeneous agents. Simula-tion results demonstrate that the proposed algorithm effectively overcomes the challenge of sparse rewards, ena-bles distributed cooperative missile interception, and significantly improves the interception success rate against adversarial anti-ship missiles.

Key words: Task Allocation, Heterogeneous-Agent Reinforcement Learning, Missile Interception, HAPPO Algorithm, Hit-rate Prediction

中图分类号: