针对协同反导拦截响应迟缓、效率低下的难题,提出一种基于异构智能体近端策略优化(HAPPO)的空海协同任务分配方法。首先,建立雷达探测跟踪、导弹动力学及毁伤效能模型,并在多智能体强化学习框架下对拦截分配问题进行形式化描述。在此基础上,为解决固定步长导致的决策冗余,设计基于事件触发的步长调整机制,实现了事件触发下的任务重构。为提升拦截任务分配策略的训练稳定性,提出了基于HAPPO的分阶段训练算法,通过渐进引入基于拦截命中率预测网络的增强奖励,提高了异构智能体的协同决策能力。仿真结果表明,所提算法有效克服了稀疏奖励的缺陷,具备空海协同任务分配能力,显著提升了对敌方反舰导弹的拦截成功率。
To address the issues of slow response and low interception efficiency in traditional task allocation methods in co-operative missile interception, this paper proposes an air-sea cooperative missile interception task allocation meth-od based on Heterogeneous-Agent Proximal Policy Optimization (HAPPO). First, models of radar detection and tracking, missile dynamics, and damage effectiveness are established, and the missile interception allocation prob-lem is formally formulated within a heterogeneous multi-agent reinforcement learning framework. On this basis, a step adjustment mechanism based on event-triggered is developed to achieve task reconfiguration upon triggering events. To enhance the training stability of the interception task allocation policy, a staged training algorithm based on HAPPO is introduced, which progressively incorporates an enhanced reward from an interception hit-rate pre-diction network, thereby enhancing the cooperative decision-making capability of heterogeneous agents. Simula-tion results demonstrate that the proposed algorithm effectively overcomes the challenge of sparse rewards, ena-bles distributed cooperative missile interception, and significantly improves the interception success rate against adversarial anti-ship missiles.