首页 >

基于改进强化学习的固定翼无人机集群事件触发编队跟踪控制

鲍忠翔1,侯懿1,于江龙1,李晓多1,李清东2,董希旺3   

  1. 1. 北京航空航天大学
    2. 北航自动化
    3. 北京航空航天大学飞行器控制一体化技术重点实验室
  • 收稿日期:2025-12-02 修回日期:2026-09-04 出版日期:2026-09-10 发布日期:2026-09-10
  • 通讯作者: 鲍忠翔

Event triggered formation tracking control of fixed wing UAV swarm Based on Improved Reinforcement Learning

  • Received:2025-12-02 Revised:2026-09-04 Online:2026-09-10 Published:2026-09-10
  • Contact: Zhongxiang bao

摘要: 针对固定翼无人机在通信资源受限、存在建模不确定性和干扰的情况下的编队问题进行了研究,提出了一种新型的结合强化学习的分层领航-跟随的编队控制方案。首先,设计了基于动态事件触发机制的分布式状态观测器,旨在节约通信资源的同时利用局部信息实现对领导者状态的精确估计,确保不会发生芝诺现象。随后,为了能补偿系统不确定性和风扰,利用RBF神经网络构建了轻量级Actor-Critic算法,并依据固定时间理论设计滑模面与收敛律,形成了基于改进强化学习的跟随者无人机集群编队跟踪滑模控制器。通过李雅普诺夫稳定性理论严格证明了观测器与控制器的收敛性。最后,通过仿真验证所提算法的有效性。

关键词: 强化学习, 固定翼无人机集群, 分布式观测器, 动态事件触发, 编队跟踪控制

Abstract: Aiming at the formation problem of fixed wing UAV under the conditions of limited communication resources, modeling uncertainty and interference, a new hierarchical pilot follow formation control scheme combined with reinforcement learn-ing is proposed. Firstly, a distributed state observer based on dynamic event triggering mechanism is designed to save communication resources and use local information to accurately estimate the leader's state, so as to ensure that Zeno phenomenon will not occur. Then, in order to compensate the system uncertainty and wind disturbance, a sliding mode tracking controller combined with reinforcement learning is proposed, and a lightweight actor critical algorithm is con-structed by using RBF neural network, and the sliding mode surface and convergence law are designed according to the fixed time theory. The convergence of the observer and controller is strictly proved by Lyapunov stability theory. Finally, the effectiveness of the proposed algorithm is verified by simulation.

Key words: reinforcement learning, fixed wing UAV swarm, distributed observer, dynamic event triggering, formation tracking control

中图分类号: