导航

Acta Aeronautica et Astronautica Sinica

Previous Articles     Next Articles

Hybrid Residual Reinforcement Learning for 3D Interception of Highly Maneuvering Targets

  

  • Received:2026-05-21 Revised:2026-09-04 Online:2026-09-10 Published:2026-09-10
  • Contact: Ying NAN

Abstract: To address the three-dimensional terminal guidance problem against highly maneuvering targets, this paper proposes a hybrid residual reinforcement learning guidance method that combines physics-based prior embedded in proportional navigation guidance (PNG) with proximal policy optimization. The proposed method takes proportional navigation guidance as the nominal baseline, while the reinforcement learning agent outputs bounded multiplicative residual gains to dynamically modify the baseline guidance command. When the neural network output is zero, the guidance law reduces to conventional PNG, thereby preserving the PNG command structure and enabling reversion to PNG under zero network output, while reducing the exploration burden of end-to-end reinforcement learning. To improve the agent’s perception of target maneuvering trends and line-of-sight rate variations, a temporal frame-stacking state representation is constructed. A dynamic entropy-coefficient decay strategy is adopted during training to regulate exploration, and first-order autopilot lag and overload saturation constraints are included in the simulation environment to improve consistency with engineering conditions. Simulation results show that, under both weaving and dynamic multi-stage random maneuver scenarios, the proposed method effectively reduces the miss distance compared with conventional proportional navigation guidance and augmented proportional navigation guidance. Although the proposed method does not achieve the lowest miss distance obtained by pure reinforcement learning guidance in the fixed maneuver case, it suppresses long-tail miss-distance risk in dynamic random maneuver scenarios and exhibits better worst-case stability. Ablation results indicate that temporal frame stacking is the main contributor to the improvement of guidance performance and robustness, while dynamic entropy-coefficient decay mainly regulates the exploration process during training. In contrast, terminal residual attenuation did not provide consistent performance gains and slightly weakened the terminal residual correction capability under the task setting of this paper; therefore, the final method adopts a hybrid residual guidance structure without terminal residual attenuation.

Key words: Deep Reinforcement Learning, Residual Control, Terminal Guidance, Proportional Navigation, Highly Maneuvering Targets

CLC Number: