首页 >

基于通道教师策略约束强化学习的无人机容错策略(飞行器协同作战技术专栏)

刘贞报,王开,贾真,王潇,颜伟俊   

  1. 西北工业大学
  • 收稿日期:2026-01-13 修回日期:2026-09-06 出版日期:2026-09-10 发布日期:2026-09-10
  • 通讯作者: 王开
  • 基金资助:
    国家自然科学基金;中国博士后科学基金;CPSF博士后奖学金计划;陕西省博士后研究资助计划;中国航空科学基金

Channel-Wise Teacher-Policy Constrained Reinforcement Learning for UAV Fault Tolerance Strategy

  • Received:2026-01-13 Revised:2026-09-06 Online:2026-09-10 Published:2026-09-10
  • Contact: Kai Wang

摘要: 渐进性电机退化故障对多旋翼无人机飞行安全构成严重威胁,为此,本文提出一种基于通道教师策略约束强化学习(CTCRL)的无人机容错策略。首先建立电机渐进退化故障模型以模拟故障动态响应;进而设计分通道控制分配算法,有效实现故障电机隔离与任务重构;在此基础上引入教师策略约束机制,以进一步增强故障下的控制重构能力。实验结果表明,在所设轻度故障下,CTCRL算法跟踪误差峰值为0.254 m,优于所对比的非坠毁算法;在所设中度故障下,CTCRL算法使故障电机平均温度相较于无温度约束对比算法降低最高达5.95 ℃;在所设重度故障下,CTCRL算法以199.76 J的最低触地动能坠落,仅为自由落体动能的14.6%,且在轨迹可控性上最优。所设硬件在环实验进一步验证该算法能够有效保证故障无人机在渐进性电机故障下的可靠性与鲁棒性,为集群编队任务持续能力提供底层技术支撑。

关键词: 容错控制, 控制分配, 多旋翼无人机, 强化学习, 电机故障

Abstract: Progressive motor degradation faults pose a serious threat to the flight safety of multi-rotor UAVs. To address this issue, this paper proposes a fault-tolerant control strategy based on Channel-wise Teacher-Policy Constrained Reinforcement Learning (CTCRL). First, a progressive motor degradation and temperature rise model is established to simulate the dynamic response of faults. Then, a channel-wise control allocation algorithm is designed to achieve effective isolation of faulty motors and mission reconfiguration. On this basis, a teacher policy constraint mechanism is introduced to further enhance the control reconfiguration capability under faults. Experimental results show that under the prescribed mild fault conditions, the CTCRL algorithm achieves a peak tracking error of 0.254 m, outperforming the compared non-crash algorithms. Under the prescribed moderate fault conditions, CTCRL reduces the average temperature of faulty motors by up to 5.95℃ compared to the contrast algorithm without temperature constraints. Under the prescribed severe fault conditions, CTCRL lands with the lowest impact kinetic energy of 199.76 J, which is only 14.6% of the free-fall kinetic energy, and exhibits the best trajectory controllability. The prescribed Hardware-in-the-loop experiments further verify that the proposed algorithm can effectively ensure the reliability and robustness of faulty UAV under progressive motor degradation, thereby providing underlying technical support for the mission sustainability of swarm formations.

Key words: Fault-Tolerant Control, Control allocation, Multi-rotor UAV, Reinforcement learning, Motor faults

中图分类号: