邵典1,2,3(
), 唐矗1,2,3, 昌敏1,2,3, 刘黎可4, 王雨乐1,2,3, 李浩2,5, 白俊强1,2,3
收稿日期:2025-11-27
修回日期:2026-01-04
接受日期:2026-02-04
出版日期:2026-02-28
发布日期:2026-02-27
通讯作者:
邵典
E-mail:shaodian@nwpu.edu.cn
基金资助:
Dian SHAO1,2,3(
), Chu TANG1,2,3, Min CHANG1,2,3, Like LIU4, Yule WANG1,2,3, Hao LI2,5, Junqiang BAI1,2,3
Received:2025-11-27
Revised:2026-01-04
Accepted:2026-02-04
Online:2026-02-28
Published:2026-02-27
Contact:
Dian SHAO
E-mail:shaodian@nwpu.edu.cn
Supported by:摘要:
以大语言模型、视觉语言模型与视觉基础模型为代表的大型基础模型(LFMs),正推动无人飞行器(UAVs)智能化的新一轮演进。围绕该趋势,首先对相关模型的关键特性与通用能力进行了归纳,总结其驱动的主流具身架构,并比较不同架构在无人飞行器高动态、强约束场景下的适配性。其次,阐明了各类大型基础模型如何通过开放环境理解、任务级语义规划、具身推理控制及多模态交互等方式,重塑无人飞行器感知、规划、控制与交互四大核心功能要素。进一步聚焦大型基础模型驱动的高阶认知功能,探讨了推理、记忆、反思与想象在应对无人飞行器复杂场景下的作用机制、实现途径、技术局限及评测范式。总结了大型基础模型在视觉-语言导航、主动目标搜索、语义物流配送及集群智能协同等4类典型决策任务中的赋能模式与前沿进展。最后,深入讨论了安全风险与防护机制、工程落地与端侧部署等核心挑战及应对策略,并从高效基础智能构建、感知-认知跨越及泛在智能产业融合等方面展望了未来发展方向。
中图分类号:
邵典, 唐矗, 昌敏, 刘黎可, 王雨乐, 李浩, 白俊强. 大型基础模型赋能下无人飞行器智能化:进展、应用与展望[J]. 航空学报, 2026, 47(15): 333148.
Dian SHAO, Chu TANG, Min CHANG, Like LIU, Yule WANG, Hao LI, Junqiang BAI. Large foundation models empowering ummannod aerial vehicle intelligence: Progress, applications and perspectives[J]. Acta Aeronautica et Astronautica Sinica, 2026, 47(15): 333148.
表1
4类典型UAV决策任务的系统性综述:涵盖基准、评估指标、代表工作、LFM赋能范式及部署验证的对比分析
| 任务类型 | 代表基准和指标 | 方法 | 相关LFM | LFM赋能范式 | 模型部署 平台 | 实验 类型 |
|---|---|---|---|---|---|---|
视觉-语言 导航 (4.1节) | 代表基准: AerialVLN、CityNav 常用指标: NE、SR、OSR、SPL、sDTW | STMR | GPT-4o (LLM) | 指令+环境表征到动作推理 | 云 (API) | 仿真+ 实机 |
Grounding DINO (VFM) | 开放词汇目标检测 | 机载平台 | ||||
Tokenize Anything (VFM) | 精细化语义分割 | |||||
| CityNavAgent | GPT-4V (VLM) | 生成场景物体描述; 多模态融合输出高层子目标 | 云 (API) | 仿真 | ||
Grounding DINO (VFM) | 开放词汇目标检测 | 本地服务器 | ||||
Segment Anything (VFM) | 精细化语义分割 | |||||
| GeoNav | GPT-3.5-turbo (LLM) | 指令解析提取目标属性 | 云 (API) | 仿真 | ||
| GPT-4o (VLM) | 构建分层场景图并导航 | |||||
Grounding DINO (VFM) | 开放词汇目标检测 | 本地服务器 | ||||
Segment Anything (VFM) | 精细化语义分割 | |||||
主动 目标搜索 (4.2节) | 代表基准: OpenUAV、CityAVOS 常用指标:NE、SR、MSS、OSR、SPL | PRPSearcher | GPT-4o (VLM) | 构建3D认知地图指导探索 | 云 (API) | 仿真 |
| Grounded-SAM (VFM) | 精细化语义分割 | 本地服务器 | ||||
| UAV-VLRR | GPT-4o (LLM) | 指令解析提取目标类别 | 云 (API) | 仿真+ 实机 | ||
| Molmo-7B (VLM) | 像素级目标与障碍定位 | 本地服务器 | ||||
语义 物流配送 (4.3节) | 代表基准: VLD 常用指标: SR、SPL、 Average Steps | LogisticsVLN | DeepSeek-R1-Distill-Qwen-14B (LLM) | CoT解析提取楼层与标志物 | 未提及 (适用边/端侧部署) | 仿真 |
| Qwen2-VL-7B (VLM) | 楼层估计/目标检测/探索 | |||||
集群 智能协同 (4.4节) | 代表基准: 模拟或现实环境 常用指标:NASA-TLX、编队形状识别准确率、用户问卷 | FlockGPT | GPT-4 (LLM) | 由自然语言解析编队参数 | 云 (API) | 仿真+ 实机 |
| LLM-CRF | LLaVA-1.6 (VLM) | 由视觉输入解析结构化语义 | 本地服务器 | 仿真 | ||
| Qwen-14B-Chat (LLM) | 多模态融合生成包含摘要、推理链和可执行代码的可执行任务包 |
| [1] | BARNHART R K, MARSHALL D M, SHAPPEE E. Introduction to unmanned aircraft systems (3rd ed.)[M]. Boca Raton: CRC Press, 2021: 56-70. |
| [2] | ZHANG W, GAO C, YU S, et al. CityNavAgent: Aerial vision-and-language navigation with hierarchical semantic planning and global memory[DB/OL]. arXiv Preprint: 2505.05622; 2025. |
| [3] | RAVICHANDRAN Z, MURALI V, TZES M, et al. SPINE: Online semantic planning for missions with incomplete natural language specifications in unstructured environments[C]∥2025 IEEE International Conference on Robotics and Automation (ICRA). Piscataway: IEEE Press, 2025. |
| [4] | WANG X, YANG D, LIAO Y, et al. UAV-flow Colosseo: A real-world benchmark for flying-on-a-word UAV imitation learning[DB/OL]. arXiv Preprint: 2505.15725; 2025. |
| [5] | WANG Z, CHEN J, ZHENG X, et al. Hi AirStar, guide me to the badminton court[C]∥Proceedings of the 33rd ACM International Conference on Multimedia. New York: ACM, 2025. |
| [6] | LU P, PENG B, CHENG H, et al. Chameleon: Plug-and-play compositional reasoning with large language models[J]. Advances in Neural Information Processing Systems, 2023, 36: 43447-43478. |
| [7] | KE F, CAI Z, JAHANGARD S, et al. HYDRA: A hyper agent for dynamic compositional visual reasoning[C]∥LEONARDIS A, RICCI E, ROTH S, et al. Computer vision-ECCV 2024. Cham: Springer Nature Switzerland, 2025. |
| [8] | SHAH D J, RUSHTON P, SINGLA S, et al. Rethinking reflection in pre-training[DB/OL]. arXiv Preprint: 2504.04022; 2025. |
| [9] | YANG Z, YU X, CHEN D, et al. Machine mental imagery: Empower multimodal reasoning with latent visual tokens[DB/OL]. arXiv Preprint: 2506.17218; 2025. |
| [10] | DEVLIN J, CHANG M W, LEE K, et al. BERT: Pre-training of deep bidirectional transformers for language understanding[C]∥2019 Conference of the North American. Minneapolis: Association for Computational Linguistics, 2019. |
| [11] | RADFORD A, WU J, CHILD R, et al. Language models are unsupervised multitask learners[J]. OpenAI Blog, 2019, 1(8): 9. |
| [12] | YANG Z, DAI Z, YANG Y, et al. XLNet: Generalized autoregressive pretraining for language understanding[J]. Advances in Neural Information Processing Systems, 2019, 32. |
| [13] | RAFFEL C, SHAZEER N, ROBERTS A, et al. Exploring the limits of transfer learning with a unified text-to-text transformer[J]. Journal of Machine Learning Research, 2020, 21(140): 1-67. |
| [14] | OPENAI. Introducing GPT-5[EB/OL]. (2025-09-02)[2025-11-27]. . |
| [15] | Introducing OpenAI o 3 and o4 -mini[EB/OL]. (2025-04-16) [2025-11-27]. . |
| [16] | ANTHROPIC. Introducing Claude 4[EB/OL]. (2025-05-22)[2025-11-27]. . |
| [17] | GUO D, YANG D, ZHANG H, et al. Deepseek-r1: Incentivizing reasoning capability in LLMs via reinforcement learning[DB/OL]. arXiv Preprint: 2501.12948; 2025. |
| [18] | YANG A, LI A, YANG B, et al. Qwen3 technical report[DB/OL]. arXiv Preprint: 2505.09388; 2025. |
| [19] | BROWN T, MANN B, RYDER N, et al. Language models are few-shot learners[J]. Advances in Neural Information Processing Systems, 2020, 33: 1877-1901. |
| [20] | WEI J, WANG X, SCHUURMANS D, et al. Chain-of-thought prompting elicits reasoning in large language models[J]. Advances in Neural Information Processing Systems, 2022, 35: 24824-24837. |
| [21] | SHAZEER N, MIRHOSEINI A, MAZIARZ K, et al. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer[DB/OL]. arXiv Preprint: 1701.06538; 2017. |
| [22] | TAN H, BANSAL M. LXMERT: Learning cross-modality encoder representations from transformers[DB/OL]. arXiv Preprint: 1908.07490; 2019. |
| [23] | LI X, YIN X, LI C, et al. Oscar: Object-semantics aligned pre-training for vision-language tasks[C]∥Computer Vision-ECCV 2020. Cham: Springer International Publishing, 2020. |
| [24] | CHEN Y C, LI L, YU L, et al. UNITER: UNiversal image-TExt representation learning[C]∥Computer Vision-ECCV 2020. Cham: Springer Cham, 2020. |
| [25] | SUN C, MYERS A, VONDRICK C, et al. VideoBERT: A joint model for video and language representation learning[C]∥Proceedings of the IEEE/CVF international conference on computer vision. Piscataway: IEEE Press, 2019. |
| [26] | LI L H, YATSKAR M, YIN D, et al. VisualBERT: A simple and performant baseline for vision and language[DB/OL]. arXiv Preprint: 1908.03557; 2019. |
| [27] | LIU H T, LI C Y, WU Q Y, et al. Visual instruction tuning[J]. Advances in Neural Information Processing Systems, 2023, 36: 34892-34916. |
| [28] | ALAYRAC J B, DONAHUE J, LUC P, et al. Flamingo: A visual language model for few-shot learning[J]. Advances in Neural Information Processing Systems, 2022, 35: 23716-23736. |
| [29] | LI J, LI D, XIONG C, et al. BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation[C]∥International conference on machine learning. New York: PMLR, 2022. |
| [30] | RADFORD A, KIM J W, HALLACY C, et al. Learning transferable visual models from natural language supervision[C]∥International conference on machine learning. New York: PMLR, 2021. |
| [31] | DOSOVITSKIY A. An image is worth 16×16 words: Transformers for image recognition at scale[DB/OL]. arXiv Preprint: 2010.11929; 2020. |
| [32] | OPENAI. GPT-4V(ision) system card[EB/OL]. (2025-08-25) [2025-11-27]. . |
| [33] | HURST A, LERER A, GOUCHER A P, et al. Gpt-4o system card[DB/OL]. arXiv Preprint: 2410.21276; 2024. |
| [34] | ANTHROPIC. Introducing the next generation of Claude[EB/OL]. (2025-08-27)[2025-11-27]. . |
| [35] | COMANICI G, BIEBER E, SCHAEKERMANN M, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities[DB/OL]. arXiv Preprint: 2507.06261; 2025. |
| [36] | BAI J, BAI S, YANG S, et al. Qwen-VL: A versatile vision-language model for understanding, localization, text reading, and beyond[DB/OL]. arXiv Preprint: 2308.12966; 2023. |
| [37] | CHEN Z, WU J, WANG W, et al. InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks[C]∥Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. Piscataway: IEEE Press, 2024. |
| [38] | WANG W, LV Q, YU W, et al. CogVLM: Visual expert for pretrained language models[J]. Advances in Neural Information Processing Systems, 2024, 37: 121475-121499. |
| [39] | CARON M, TOUVRON H, MISRA I, et al. Emerging properties in self-supervised vision transformers[C]∥Proceedings of the IEEE/CVF international conference on computer vision. Piscataway: IEEE Press, 2021. |
| [40] | CARON M, TOUVRON H, MISRA I, et al. Emerging properties in self-supervised vision transformers[C]∥Proceedings of the IEEE/CVF international conference on computer vision. Piscataway: IEEE Press, 2021. |
| [41] | OQUAB M, DARCET T, MOUTAKANNI T, et al. Dinov2: Learning robust visual features without supervision[DB/OL]. arXiv Preprint: 2304.07193; 2023. |
| [42] | KIRILLOV A, MINTUN E, RAVI N, et al. Segment anything[C]∥Proceedings of the IEEE/CVF international conference on computer vision. Piscataway: IEEE Press, 2023. |
| [43] | LI L H, ZHANG P, ZHANG H, et al. Grounded language-image pre-training[C]∥Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. Piscataway: IEEE Press, 2022. |
| [44] | CHENG T, SONG L, GE Y, et al. Yolo-world: Real-time open-vocabulary object detection[C]∥Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. Piscataway: IEEE Press, 2024. |
| [45] | LIU S, ZENG Z, REN T, et al. Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection[C]∥Computer Vision-ECCV 2024. Cham: Springer Cham, 2025. |
| [46] | LÜDDECKE T, ECKER A. Image segmentation using text and image prompts[C]∥Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. Piscataway: IEEE Press, 2022. |
| [47] | YUAN H, LI X, ZHOU C, et al. Open-vocabulary SAM: Segment and recognize twenty-thousand classes interactively[C]∥Computer Vision-ECCV 2024. Cham: Springer Cham, 2025. |
| [48] | WANG X, ZHANG X, CAO Y, et al. SegGPT: towards segmenting everything in context[C]∥Proceedings of the IEEE/CVF International Conference on Computer Vision. Piscataway: IEEE Press, 2023. |
| [49] | ZOU X, YANG J, ZHANG H, et al. Segment everything everywhere all at once[J]. Advances in Neural Information Processing Systems, 2023, 36: 19769-19782. |
| [50] | BHAT S F, BIRKL R, WOFK D, et al. ZoeDepth: Zero-shot transfer by combining relative and metric depth[DB/OL]. arXiv Preprint: 2302.12288; 2023. |
| [51] | YANG L, KANG B, HUANG Z, et al. Depth anything: Unleashing the power of large-scale unlabeled data[C]∥Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. Piscataway: IEEE Press, 2024. |
| [52] | BOCHKOVSKII A, DELAUNOY A, GERMAIN H, et al. Depth pro: Sharp monocular metric depth in less than a second[DB/OL]. arXiv Preprint: 2410.02073; 2024. |
| [53] | ZITKOVICH B, YU T, XU S, et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control[C]∥Conference on Robot Learning. New York: PMLR, 2023. |
| [54] | KIM M J, PERTSCH K, KARAMCHETI S, et al. Openvla: An open-source vision-language-action model[DB/OL]. arXiv Preprint: 2406.09246; 2024. |
| [55] | WEN J, ZHU Y, LI J, et al. TinyVLA: Toward fast, data-efficient vision-language-action models for robotic manipulation[J]. IEEE Robotics and Automation Letters, 2025, 10(4): 3988-3995. |
| [56] | MEI A, ZHU G N, ZHANG H, et al. ReplanVLM: Replanning robotic tasks with visual language models[J]. IEEE Robotics and Automation Letters, 2024, 9(11): 10201-10208. |
| [57] | SKRETA M, ZHOU Z, YUAN J L, et al. Replan: Robotic replanning with perception and language models[DB/OL]. arXiv Preprint: 2401.04157; 2024. |
| [58] | AO S, SALIM F D, KHAN S. EMAC+: Embodied multimodal agent for collaborative planning with VLM+ LLM[DB/OL]. arXiv Preprint: 2505.19905; 2025. |
| [59] | ZHAO Q, LU Y, KIM M J, et al. Cot-vla: Visual chain-of-thought reasoning for vision-language-action models[C]∥Proceedings of the Computer Vision and Pattern Recognition Conference. Piscataway: IEEE Press, 2025. |
| [60] | SUN Q, HONG P, PALA T D, et al. Emma-X: An embodied multimodal action model with grounded chain of thought and look-ahead spatial reasoning[C]∥63rd Annual Meeting of the Association for Computational Linguistics Vienna: Association for Computational Linguistics, 2025. |
| [61] | LIN F, NAI R, HU Y, et al. OneTwoVLA: A unified vision-language-action model with adaptive reasoning[DB/OL]. arXiv Preprint: 2505.11917; 2025. |
| [62] | SERPIVA V, LYKOV A, MYSHLYAEV A, et al. RaceVLA: VLA-based racing drone navigation with human-like behaviour[DB/OL]. arXiv Preprint: 2503.02572; 2025. |
| [63] | HE K, ZHANG X, REN S, et al. Deep residual learning for image recognition[C]∥Proceedings of the IEEE conference on computer vision and pattern recognition. Piscataway: IEEE Press, 2016. |
| [64] | HUSSAIN M. YOLO-v1 to YOLO-v8, the rise of YOLO and its complementary nature toward digital manufacturing and industrial defect detection[J]. Machines, 2023, 11(7): 677. |
| [65] | REN S, HE K, GIRSHICK R, et al. Faster R-CNN: Towards real-time object detection with region proposal networks[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2016, 39(6): 1137-1149. |
| [66] | ZHANG X, TIAN Y, LIN F, et al. LogisticsVLN: Vision-language navigation for low-altitude terminal delivery based on agentic UAVs[DB/OL]. arXiv Preprint: 2505.03460; 2025. |
| [67] | HART P E, NILSSON N J, RAPHAEL B. A formal basis for the heuristic determination of minimum cost paths[J]. IEEE transactions on Systems Science and Cybernetics, 1968, 4(2): 100-107. |
| [68] | LAVALLE S. Rapidly-exploring random trees: A new tool for path planning[R]. Research Report 9811, 1998. |
| [69] | SHAH D, OSIŃSKI B, LEVINE S, et al. Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action[C]∥Conference on robot learning. New York: PMLR, 2023. |
| [70] | AHN M, BROHAN A, BROWN N, et al. Do as I can, not as I say: Grounding language in robotic affordances[DB/OL]. arXiv Preprint: 2204.01691; 2022. |
| [71] | XIAO J, TSAO C W, ZHANG Y, et al. FM-planner: Foundation model guided path planning for autonomous drone navigation[DB/OL]. arXiv Preprint: 2505.20783; 2025. |
| [72] | BAI J, BAI S, CHU Y, et al. Qwen technical report[DB/OL]. arXiv Preprint: 2309.16609; 2023. |
| [73] | TEAM G, MESNARD T, HARDIN C, et al. Gemma: Open models based on Gemini research and technology[DB/OL]. arXiv Preprint: 2403.08295; 2024. |
| [74] | DUBEY A, JAUHRI A, PANDEY A, et al. The llama 3 herd of models[DB/OL]. arXiv Preprint: 2407.21783; 2024. |
| [75] | WATKINS C J, DAYAN P. Q-learning[J]. Machine learning, 1992, 8(3-4): 279-292. |
| [76] | CHEN G, YU X, LING N, et al. TypeFly: Flying drones with large language model[DB/OL]. arXiv Preprint: 2312.14950; 2023. |
| [77] | TAGLIABUE A, KONDO K, ZHAO T, et al. REAL: Resilience and adaptation using large language models on autonomous aerial robots[C]∥2024 IEEE 63rd Conference on Decision and Control (CDC). Piscataway: IEEE Press, 2024. |
| [78] | WANG W, LI Y, JIAO L, et al. GSCE: A prompt framework with enhanced reasoning for reliable LLM-driven drone control[C]∥2025 International Conference on Unmanned Aircraft Systems (ICUAS). Piscataway: IEEE Press, 2025. |
| [79] | YOO M, NA Y, SONG H, et al. Motion estimation and hand gesture recognition-based human-UAV interaction approach in real time[J]. Sensors, 2022, 22(7): 2513. |
| [80] | 万开方, 吴志林, 武韫晖, 等. 拒止环境下基于深度强化学习的多无人机协同定位[J]. 航空学报, 2025, 46(8): 331024. |
| WAN K F, WU Z L, WU Y H, et al. Multi-UAV cooperative localization based on deep reinforcement learning in denied environments[J]. Acta Aeronautica et Astronautica Sinica, 2025, 46(8): 331024 (in Chinese). | |
| [81] | CLADERA F, RAVICHANDRAN Z, HUGHES J, et al. Air-ground collaboration for language-specified missions in unknown environments[DB/OL]. arXiv Preprint: 2505.09108; 2025. |
| [82] | RAVICHANDRAN Z, CLADERA F, HUGHES J, et al. Deploying foundation model-enabled air and ground robots in the field: Challenges and opportunities[DB/OL]. arXiv Preprint: 2505.09477; 2025. |
| [83] | LYKOV A, KARAF S, MARTYNOV M, et al. FlockGPT: Guiding UAV flocking with linguistic orchestration[C]∥2024 IEEE International Symposium on Mixed and Augmented Reality Adjunct (ISMAR-Adjunct). Piscataway: IEEE Press, 2024. |
| [84] | DHARMALINGAM B, MUKHERJEE R, PIGGOTT B, et al. Aero-LLM: A distributed framework for secure UAV communication and intelligent decision-making[DB/OL]. arXiv Preprint: 2502.05220; 2025. |
| [85] | DAI S, MA Z, LUO Z, et al. MM-UAVBench: How well do multimodal large language models see, think, and plan in low-altitude UAV scenarios?[DB/OL]. arXiv Preprint: 2512.23219; 2025. |
| [86] | ZHU Z, XUE Y, CHEN X, et al. Large language models can learn rules[DB/OL]. arXiv Preprint: 2310.07064; 2023. |
| [87] | HAO S B, GU Y, MA H D, et al. Reasoning with language model is planning with world model[DB/OL]. arXiv Preprint: 2305.14992; 2023. |
| [88] | MOWER C E, WAN Y, YU H, et al. ROS-LLM: A ros framework for embodied ai with task feedback and structured reasoning[DB/OL]. arXiv Preprint: 2406.19741; 2024. |
| [89] | ZHANG W, WANG M, LIU G, et al. Embodied-reasoner: Synergizing visual search, reasoning, and action for embodied interactive tasks[DB/OL]. arXiv Preprint: 2503.21696; 2025. |
| [90] | CAI Z, CARDENAS C R, LEO K, et al. NEUSIS: A compositional neuro-symbolic framework for autonomous perception, reasoning, and planning in complex UAV search missions[J]. IEEE Robotics and Automation Letters, 2025, 10(9): 9502-9509. |
| [91] | GAO Y, WANG Z, JING L, et al. Aerial vision-and-language navigation via semantic-topo-metric representation guided LLM reasoning[DB/OL]. arXiv Preprint: 2410.08500; 2024. |
| [92] | ZHAO J, LIN X. General-purpose aerial intelligent agents empowered by large language models[DB/OL]. arXiv Preprint: 2503.08302; 2025. |
| [93] | ZHAO B, WANG Z, FANG J, et al. Embodied-R: Collaborative framework for activating embodied spatial reasoning in foundation models via reinforcement learning[C]∥Proceedings of the 33rd ACM International Conference on Multimedia. New York: ACM, 2025. |
| [94] | WANG G, XIE Y, JIANG Y, et al. Voyager: An open-ended embodied agent with large language models[DB/OL]. arXiv Preprint: 2305.16291; 2023. |
| [95] | WANG Z, LI X, YANG J, et al. Gridmm: Grid memory map for vision-and-language navigation[C]∥Proceedings of the IEEE/CVF International conference on computer vision. Piscataway: IEEE Press, 2023. |
| [96] | LI H, WANG Z, YANG X, et al. Memonav: Working memory model for visual navigation[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2024. |
| [97] | ZHENG Q, LIU D, WANG C, et al. Esceme: Vision-and-language navigation with episodic scene memory[J]. International Journal of Computer Vision, 2025, 133(1): 254-274. |
| [98] | ZENG Q, YANG Q, DONG S, et al. Perceive, reflect, and plan: Designing LLM agent for goal-directed city navigation without instructions[DB/OL]. arXiv Preprint: 2408.04168; 2024. |
| [99] | LINGO R, ARROYO M, CHHAJER R. Enhancing LLM problem solving with reap: Reflection, explicit problem deconstruction, and advanced prompting[DB/OL]. arXiv Preprint: 2409.09415; 2024. |
| [100] | YANG Z, LIU J, CHEN P, et al. RILA: Reflective and imaginative language agent for zero-shot semantic audio-visual navigation[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2024. |
| [101] | YU Z, LONG Y, YANG Z, et al. CorrectNav: Self-correction Flywheel empowers vision-language-action navigation model[DB/OL]. arXiv Preprint: 2508.10416; 2025. |
| [102] | HUANG Y, WU M, LI R, et al. VISTA: Generative visual imagination for vision-and-language navigation[DB/OL]. arXiv Preprint: 2505.07868; 2025. |
| [103] | PERINCHERRY A, KRANTZ J, LEE S. Do visual imaginations improve vision-and-language navigation agents?[C]∥Proceedings of the Computer Vision and Pattern Recognition Conference. Piscataway: IEEE Press, 2025. |
| [104] | ANDERSON P, WU Q, TENEY D, et al. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments[C]∥Proceedings of the IEEE conference on computer vision and pattern recognition. Piscataway: IEEE Press, 2018. |
| [105] | KU A, ANDERSON P, PATEL R, et al. Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding[DB/OL]. arXiv Preprint: 2010.07954; 2020. |
| [106] | CHEN H, SUHR A, MISRA D, et al. Touchdown: Natural language navigation and spatial reasoning in visual street environments[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE Press, 2019. |
| [107] | LEE J, MIYANISHI T, KURITA S, et al. CityNav: Language-goal aerial navigation dataset with geographic information[DB/OL]. arXiv Preprint: 2406.14240; 2024. |
| [108] | DORBALA V S, SIGURDSSON G, PIRAMUTHU R, et al. CLIP-Nav: Using clip for zero-shot vision-and-language navigation[DB/OL]. arXiv Preprint: 2211.16649; 2022. |
| [109] | DORBALA V S, MULLEN JR J F, MANOCHA D. Can an embodied agent find your “cat-shaped mug”? LLM-based zero-shot object navigation[DB/OL]. arXiv Preprint: 2303.03480; 2023. |
| [110] | ZHOU G, HONG Y, WU Q. NavGPT: Explicit reasoning in vision-and-language navigation with large language models[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2024, 38(7): 7641-7649. |
| [111] | LONG Y, LI X, CAI W, et al. Discuss before moving: Visual language navigation via multi-expert discussions[C]∥2024 IEEE International Conference on Robotics and Automation (ICRA). Piscataway: IEEE Press, 2024. |
| [112] | LONG Y, CAI W, WANG H, et al. InstructNav: Zero-shot system for generic instruction navigation in unexplored environment[DB/OL]. arXiv Preprint: 2406.04882; 2024. |
| [113] | ZHOU G, HONG Y, WANG Z, et al. NavGPT-2: Unleashing navigational reasoning capability for large vision-language models[C]∥Computer Vision-ECCV 2024. Cham: Springer Cham, 2025. |
| [114] | SU Y, AN D, CHEN K, et al. Learning fine-grained alignment for aerial vision-dialog navigation[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2025, 39(7): 7060-7068. |
| [115] | XU H, HU Y, GAO C, et al. GeoNav: Empowering MLLMs with explicit geospatial reasoning abilities for language-goal aerial navigation[DB/OL]. arXiv Preprint: 2504.09587; 2025. |
| [116] | ZHAO G, LI G, PAN J, et al. Aerial vision-and-language navigation with grid-based view selection and map construction[DB/OL]. arXiv Preprint: 2503.11091; 2025. |
| [117] | LIN H, YANG X, WEN G, et al. Fast UAV object-searching in large-scale and complex environments[J]. IEEE Transactions on Cybernetics, 2025, 55(6): 2993-3004. |
| [118] | CHANGOLUISA CAIZA I D, MILAS A, MONTES GROVA M A, et al. Autonomous exploration of unknown 3D environments using a frontier-based collector strategy[C]∥2024 IEEE International Conference on Robotics and Automation (ICRA). Piscataway: IEEE Press, 2024. |
| [119] | PAPAIOANNOU S, KOLIOS P, THEOCHARIDES T, et al. 3D trajectory planning for UAV-based search missions: An integrated assessment and search planning approach[C]∥2021 International Conference on Unmanned Aircraft Systems (ICUAS). Piscataway: IEEE Press, 2021. |
| [120] | BARTOLOMEI L, TEIXEIRA L, CHLI M. Fast multi-UAV decentralized exploration of forests[J]. IEEE Robotics and Automation Letters, 2023, 8(9): 5576-5583. |
| [121] | SANDINO J, CACCETTA P A, SANDERSON C, et al. Reducing object detection uncertainty from RGB and thermal data for UAV outdoor surveillance[C]∥2022 IEEE Aerospace Conference (AERO). Piscataway: IEEE Press, 2022. |
| [122] | GUPTA A, BESSONOV D, LI P. A decision-theoretic approach to detection-based target search with a UAV[C]∥2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). Piscataway: IEEE Press, 2017. |
| [123] | XU K, ZHENG L, WEI M, et al. VRExplorer: An efficient view-region based autonomous exploration method in unknown environments for UAV[C]∥2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). Piscataway: IEEE Press, 2024. |
| [124] | YAN P, JIA T, BAI C C. Searching and tracking an unknown number of targets: A learning-based method enhanced with maps merging[J]. Sensors, 2021, 21(4): 1076. |
| [125] | YAQOOT Y, MUSTAFA M A, SAUTENKOV O, et al. UAV-VLRR Vision-language informed nmpc for rapid response in UAV search and rescue[DB/OL]. arXiv Preprint: 2503.02465; 2025. |
| [126] | JI Y, ZHU Z, ZHAO Y, et al. Towards autonomous UAV visual object search in city space: Benchmark and agentic methodology[DB/OL]. arXiv Preprint: 2505.08765; 2025. |
| [127] | TIAN Y, LIN F, ZHANG X, et al. LogisticsVISTA: 3D terminal delivery services with UAVs, UGVs and USVs based on foundation models and scenarios engineering[C]∥2024 IEEE International Conference on Service Operations and Logistics, and Informatics (SOLI). Piscataway: IEEE Press, 2024. |
| [128] | 于彦鹏, 余墨多, 汤奇荣, 等. 面向城市应急物资配送的多无人机协同路径规划算法[J]. 控制与决策, 2025, 40(4): 1098-1106. |
| YU Y P, YU M D, TANG Q R, et al. Multi-UAV cooperative path planning algorithm for urban emergency logistics distribution [J]. Control and Decision, 2025, 40(4): 1098-1106 (in Chinese). | |
| [129] | 张启钱, 许卫卫, 张洪海, 等. 复杂低空物流无人机路径规划[J]. 北京航空航天大学学报, 2020, 46(7): 1275-1286. |
| ZHANG Q Q, XU W W, ZHANG H H, et al. Path planning for low-altitude logistics UAVs in complex environments [J]. Journal of Beijing University of Aeronautics and Astronautics, 2020, 46(7): 1275-1286 (in Chinese). | |
| [130] | YAO Y, LUO S, ZHAO H, et al. Can LLM substitute human labeling? A case study of fine-grained chinese address entity recognition dataset for UAV delivery[C]∥Companion Proceedings of the ACM Web Conference 2024.New York: ACM, 2024. |
| [131] | ZHOU B, XU H, SHEN S. Racer: Rapid collaborative exploration with a decentralized multi-UAV system[J]. IEEE Transactions on Robotics, 2023, 39(3): 1816-1835. |
| [132] | PAPYAN N, KULHANDJIAN M, KULHANDJIAN H, et al. AI-based drone assisted human rescue in disaster environments: Challenges and opportunities[J]. Pattern Recognition and Image Analysis, 2024, 34(1): 169-186. |
| [133] | PUEYO P, MONTIJANO E, MURILLO A C, et al. CLIPSwarm: Generating drone shows from text prompts with vision-language models[C]∥2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). Piscataway: IEEE Press, 2024. |
| [134] | ZHANG T, TIAN Y, LIN F, et al. CoordField: Coordination field for agentic UAV task allocation in low-altitude urban scenarios[DB/OL]. arXiv Preprint: 2505.00091; 2025. |
| [135] | LIN F, ZHANG T, NI Q, et al. Talk less, fly lighter: Autonomous semantic compression for UAV swarm communication via LLMs[DB/OL]. arXiv Preprint: 2508.12043; 2025. |
| [136] | JI K, HU X, ZHANG X, et al. An LLM-based framework for human-swarm teaming cognition in disaster search and rescue[DB/OL]. arXiv Preprint: 2511.04042; 2025. |
| [137] | KRINNER M, ROMERO A, BAUERSFELD L, et al. MPCC++: Model predictive contouring control for time-optimal flight with safety constraints[DB/OL]. arXiv Preprint: 2403.17551; 2024. |
| [138] | BUI D N, KHUAT T H, PHUNG M D, et al. Model predictive control for optimal motion planning of unmanned aerial vehicles[C]∥2024 International Conference on Control, Robotics and Informatics (ICCRI). Piscataway: IEEE Press, 2024. |
| [139] | BOUCHEKIR R, GUZMAN M, COOK A, et al. Formal verification for safe AI-based flight planning for UAVs*[C]∥2023 53rd Annual IEEE/IFIP International Conference on Dependable Systems and Networks Workshops (DSN-W). Piscataway: IEEE Press, 2023. |
| [140] | TANG Y C, CHEN P Y, HO T Y. Defining and evaluating physical safety for large language models[DB/OL]. arXiv Preprint: 2411.02317; 2024. |
| [141] | SUN Y, CHEN Y, WU P, et al. DRL: Dynamic rebalance learning for adversarial robustness of UAV with long-tailed distribution[J]. Computer Communications, 2023, 205: 14-23. |
| [142] | REN Y, ZHU F, LU G, et al. Safety-assured high-speed navigation for MAVs[J]. Science Robotics, 2025, 10(98): eado6187. |
| [143] | BADAR A U R, MAHMOOD D, IQBAL A, et al. DeepSpoofNet: A framework for securing UAVs against GPS spoofing attacks[J]. PeerJ Computer Science, 2025, 11: e2714. |
| [144] | NVIDIA jetson AGX orin[EB/OL]. (2025-10-28)[2025-11-27]. . |
| [145] | Atlas 200I A2[EB/OL]. (2025-11-12)[2025-11-27]. . |
| [146] | RK3588[EB/OL]. (2025-08-08)[2025-11-27]. . |
| [147] | AR9481[EB/OL]. (2021-12-20)[2025-11-27]. . |
| [148] | MCENROE P, WANG S, LIYANAGE M. A survey on the convergence of edge computing and AI for UAVs: Opportunities and challenges[J]. IEEE Internet of Things Journal, 2022, 9(17): 15435-15459. |
| [149] | XU J, LI Z, CHEN W, et al. On-device language models: A comprehensive review[DB/OL]. arXiv Preprint: 2409.00088; 2024. |
| [150] | LIN J, TANG J, TANG H, et al. Awq: Activation-aware weight quantization for on-device llm compression and acceleration[J]. Proceedings of Machine Learning and Systems, 2024, 6: 87-100. |
| [151] | GU A, DAO T. Mamba: Linear-time sequence modeling with selective state spaces[DB/OL]. arXiv Preprint: 2312.00752; 2023. |
| [1] | 傅圣杰 李明磊 金晟毅 唐佳瑞 郭一凡. 面向火星采样任务的异源图像采样区高精度重建研究-“深空探测前沿技术”专刊[J]. 航空学报, 0, (): 1-0. |
| [2] | 徐海涛 刘鹏 薛长斌 沈卫华 曹印国 张亚男 张静 檀晓萌 李娇娇 曹锴郎. 面向小行星探测的多源光轴自主标校方法-“深空探测前沿技术”专刊[J]. 航空学报, 0, (): 1-0. |
| [3] | 周雷, 谷延锋, 刘天竹. 遥感图像飞机细粒度目标检测算法[J]. 航空学报, 2026, 47(10): 632578-632578. |
| [4] | 刘涛 任侃 温世博 陈钱. 弱纹理下物理引导几何校验无人机视觉定位[J]. 航空学报, 0, (): 1-0. |
| [5] | 郝振洋, 曹尚, 张凤婷, 侯兰兰. 直升机桨毂顶置主动式作动系统设计[J]. 航空学报, 2026, 47(7): 432464-432464. |
| [6] | 赵耀, 张曦, 周荻, 李玉堂, 李思远. 考虑主动远离禁飞区的飞行器再入轨迹快速规划[J]. 航空学报, 2026, 47(6): 332509-332509. |
| [7] | 王征, 赵守智, 李昊田, 孙征, 邵静, 侯丞. 深空探测用穿冰探测器发展现状[J]. 航空学报, 2026, 47(5): 332398-332398. |
| [8] | 乔塨哲 杨盘隆 叶彤 姚世宸 付道勇. 面向单粒子效应的卫星网络弹性组网及重构方法-“空天地一体化智能网联”专刊[J]. 航空学报, 0, (): 1-0. |
| [9] | 陶飞, 张贺, 刘蔚然, 张辰源, 魏宇鹏, 易黎, 邹孝付. 空天装备数字试验验证理论与关键技术[J]. 航空学报, 2025, 46(24): 432516-432516. |
| [10] | 陈亮, 孟凡星, 王成波, 张音旋, 孟琳书. 数字孪生技术在飞行器强度设计中的发展及应用[J]. 航空学报, 2025, 46(19): 532252-532252. |
| [11] | 刘童, 赵文祥, 吉敬华. 逆变器供电永磁电机电磁振噪计算方法综述[J]. 航空学报, 2026, 47(1): 332127-332127. |
| [12] | 蒲钒, 陈志杰, 刘杨, 耿欣, 朱永文, 任柯锦. 数字低空融合运行空中交通管理技术[J]. 航空学报, 2025, 46(11): 531331-531331. |
| [13] | 刘泓麟, 王冠, 安帅斌, 马少捷, 刘凯. 基于在线辨识的高速变构飞行器强适应控制[J]. 航空学报, 2025, 46(17): 331654-331654. |
| [14] | 陈树生, 贾苜梁, 林家豪, 金世轶, 高正红, 王岳青, 马志强, 李铮, 段辰龙, 李佳伟. 生成式模型赋能飞行器技术应用研究进展与展望[J]. 航空学报, 2025, 46(10): 631194-631194. |
| [15] | 林杰, 唐志共, 钱炜祺, 王岳青, 张鹏, 徐炜遐, 刘杰. 飞行器生成式模型气动设计研究进展与展望[J]. 航空学报, 2025, 46(10): 631679-631679. |
| 阅读次数 | ||||||
|
全文 |
|
|||||
|
摘要 |
|
|||||
版权所有 © 航空学报编辑部
版权所有 © 2011航空学报杂志社
主管单位:中国科学技术协会 主办单位:中国航空学会 北京航空航天大学

