航空学报 > 2026, Vol. 47 Issue (15): 333148-333148   doi: 10.7527/S1000-6893.2026.33148

大型基础模型赋能下无人飞行器智能化:进展、应用与展望

邵典1,2,3(), 唐矗1,2,3, 昌敏1,2,3, 刘黎可4, 王雨乐1,2,3, 李浩2,5, 白俊强1,2,3   

  1. 1.西北工业大学 无人系统技术研究院,西安 710072
    2.无人飞行器技术全国重点实验室,西安 710072
    3.西北工业大学 人工智能学院,西安 710072
    4.西北工业大学 软件学院,西安 710072
    5.航空工业成都飞机设计研究所,成都 610041
  • 收稿日期:2025-11-27 修回日期:2026-01-04 接受日期:2026-02-04 出版日期:2026-02-28 发布日期:2026-02-27
  • 通讯作者: 邵典 E-mail:shaodian@nwpu.edu.cn
  • 基金资助:
    国家自然科学基金青年基金(62306239);无人飞行器技术全国重点实验室开放基金(WRFX-202424);三秦英才引进计划

Large foundation models empowering ummannod aerial vehicle intelligence: Progress, applications and perspectives

Dian SHAO1,2,3(), Chu TANG1,2,3, Min CHANG1,2,3, Like LIU4, Yule WANG1,2,3, Hao LI2,5, Junqiang BAI1,2,3   

  1. 1.Unmanned System Research Institute,Northwestern Polytechnical University,Xi’an 710072,China
    2.National Key Laboratory of Unmanned Aerial Vehicle Technology,Xi’an 710072,China
    3.School of Artificial Intelligence,Northwestern Polytechnical University,Xi’an 710072,China
    4.School of Software,Northwestern Polytechnical University,Xi’an 710072,China
    5.AVIC Chengdu Aircraft Design & Research Institute,Chengdu 610041,China
  • Received:2025-11-27 Revised:2026-01-04 Accepted:2026-02-04 Online:2026-02-28 Published:2026-02-27
  • Contact: Dian SHAO E-mail:shaodian@nwpu.edu.cn
  • Supported by:
    National Natural Science Foundation of China(62306239);Open Fund of the National Key Laboratory of Unmanned Aerial Vehicle Technology(WRFX-202424);Sanqin Talents Introduction Plan of Shaanxi Province

摘要:

以大语言模型、视觉语言模型与视觉基础模型为代表的大型基础模型(LFMs),正推动无人飞行器(UAVs)智能化的新一轮演进。围绕该趋势,首先对相关模型的关键特性与通用能力进行了归纳,总结其驱动的主流具身架构,并比较不同架构在无人飞行器高动态、强约束场景下的适配性。其次,阐明了各类大型基础模型如何通过开放环境理解、任务级语义规划、具身推理控制及多模态交互等方式,重塑无人飞行器感知、规划、控制与交互四大核心功能要素。进一步聚焦大型基础模型驱动的高阶认知功能,探讨了推理、记忆、反思与想象在应对无人飞行器复杂场景下的作用机制、实现途径、技术局限及评测范式。总结了大型基础模型在视觉-语言导航、主动目标搜索、语义物流配送及集群智能协同等4类典型决策任务中的赋能模式与前沿进展。最后,深入讨论了安全风险与防护机制、工程落地与端侧部署等核心挑战及应对策略,并从高效基础智能构建、感知-认知跨越及泛在智能产业融合等方面展望了未来发展方向。

关键词: 智能无人飞行器, 大型基础模型, 无人飞行器自主决策, 具身智能, 智能体高阶认知

Abstract:

Large Foundation Models (LFMs), represented by large language models, vision language models, and vision foundation models, are driving a new wave of intelligent evolution for Unmanned Aerial Vehicles (UAVs). Focusing on this trend, the key characteristics and general capabilities of relevant models are first summarized, followed by a categorization of the mainstream embodied architectures driven by them. A comparison is conducted regarding the adaptability and trade-offs of different architectures within the high-dynamic and strongly constrained scenarios of UAVs. Secondly, an analysis is provided on how various LFMs reshape the four core functional elements of UAVs, including perception, planning, control, and interaction, through mechanisms such as open-world understanding, task-level semantic planning, embodied reasoning control, and multi-modal interaction. Furthermore, focusing on high-level cognitive functions driven by LFMs, the mechanisms, implementation pathways, technical limitations, and evaluation paradigms of reasoning, memory, reflection, and imagination in coping with complex UAV scenarios are discussed. The empowerment patterns and frontier progress of LFMs in four typical decision-making tasks are then summarized, including vision-language navigation, active target search, semantic delivery, and swarm intelligent coordination. Finally, core challenges regarding safety risks and protection mechanisms, engineering implementation, and edge deployment are discussed, envisioning future directions in efficient foundation intelligence, the perception-to-cognition transition, and ubiquitous industrial collaboration.

Key words: intelligent UAVs, large foundation models, UAV decision-making tasks, embodied intelligence, higher-order cognition in agents

中图分类号: