跨模态注意力交互的航拍小目标智能检测算法-AI+空天科学

  • 唐鼎新 ,
  • 杨文青 ,
  • 符强 ,
  • 宣建林
展开
  • 1. 西北工业大学航空学院
    2. 西北工业大学

收稿日期: 2026-03-30

  修回日期: 2026-05-25

  网络出版日期: 2026-06-01

基金资助

715科技创新基金

Intelligent small object detection in aerial images via cross-modal attention interaction

  • TANG Ding-Xin ,
  • YANG Wen-Qing ,
  • FU Qiang ,
  • XUAN Jian-Lin
Expand

Received date: 2026-03-30

  Revised date: 2026-05-25

  Online published: 2026-06-01

摘要

在无人机航拍小目标检测任务中,如何在抑制模态间干扰的同时高效融合跨模态互补信息,是提升复杂环境下检测性能的关键问题。针对该挑战,本文提出一种双模态融合小目标检测算法PCMA-Net。首先,设计并行注意力与门控重标定模块,通过并行双路架构捕获细节特征,并利用动态门控掩码对增强特征执行二次特征精炼,抑制复杂背景伪影,提升模型在复杂背景下小目标判别能力。其次,构建跨模态注意力交互模块通过预建模响应分布聚焦显著区域,利用双向交叉注意力实现异质特征的像素级动态对齐,解决特征失配和掩没问题。最后,引入双分支注意力引导卷积,部分通道生成的动态权重引导主干路径进行特征增强,在较低计算开销下实现全局上下文信息建模,从而提升推理效率。实验结果表明,在DroneVehicle数据集上,所提方法相较于基线与主流融合方法,mAP@0.5分别提升1.9%和1.7%;在VEDAI数据集上较基线提升5.6%,并在多个类别上取得最优性能。同时,PCMA-Net实现了45.1 FPS的推理速度,满足无人机实时检测需求。

本文引用格式

唐鼎新 , 杨文青 , 符强 , 宣建林 . 跨模态注意力交互的航拍小目标智能检测算法-AI+空天科学[J]. 航空学报, 0 : 1 -0 . DOI: 10.7527/S1000-6893.2026.33627

Abstract

In unmanned aerial vehicle (UAV) aerial photography tasks, the core challenge for achieving precision detection in complex environments lies in effectively suppressing inter-modal interference while efficiently fusing cross-modal complementary features. To address this challenge, this paper proposes PCMA-Net, a dual-modal fusion algorithm for small object detection in UAV imagery. First, a parallel attention and gated recalibration module is designed. By capturing detailed features through a parallel dual-path architecture and employing dynamic gated masking for secondary feature refinement, this module suppresses complex background artifacts and enhances the model's discriminative power for small objects in complex backgrounds. Second, a cross-modal attention interaction module is constructed. By pre-modeling the response distribution to focus on salient regions, it utilizes bidirectional cross-attention to achieve pixel-level dynamic alignment of heterogeneous features, thereby resolving feature mismatch and submergence issues. Finally, we propose dual-branch attention guided convolution, which utilizes channel splitting and task-aware weight generation to achieve global context modeling with minimal overhead, thereby improving inference efficiency on embedded platforms. Experimental results on the DroneVehicle dataset demonstrate that the proposed method outperforms the baseline and state-of-the-art fusion algorithms by 1.9% and 1.7% in mAP@0.5, respectively. Furthermore, on the VEDAI dataset, PCMA-Net exceeds the baseline by 5.6% in mAP@0.5 and achieves state-of-the-art performance across most specific categories. Notably, PCMA-Net achieves a detection speed of 45.1 FPS, fully meeting the real-time requirements for UAV-based deployment.

参考文献

[1] CHENG J, DENG C, SU Y, et al.Methods and datasets on semantic segmentation for Unmanned Aerial Vehicle remote sensing images: A review[J].ISPRS Journal of Photogrammetry and Remote Sensing, 2024, 211(000):1-34
[2] LENG J, YE Y, MO M ,et al. Recent Advances for Aerial Object Detection: A Survey[J].ACM Computing Surveys, 2024, 56(12):1-36.
[3] 吴一全,童康.基于深度学习的无人机航拍图像小目标检测研究进展[J].航空学报,2025,46,(3):174-200.WU YQ,TONG K. Research advances on deep learning-based small object detection in UAV aerial images[J].Acta Aeronautica et Astronautica Sinica,2025,46(3):030848(in Chinese).
[4] TELIKANI A, SARKAR A, DU B, et al. Autonomous Aerial Vehicles-Aided Intelligent Transportation Systems: Vision, Challenges, and Opportunities[J].IEEE Commu-nications Surveys & Tutorials,2025,27,(6):3772-3819.
[5] IJAZ H, AHMAD R, AHMED R, et al. A UAV-Assisted Edge Framework for Real-Time Disaster Manage-ment[J].IEEE Transactions on Geoscience and Remote Sensing,2023,61,1-1.
[6] FANG Z X, SAVKIN, ANDREY V. Strategies for Opti-mized UAV Surveillance in Various Tasks and Scenarios: A Review[J].Drones,2024,8,(5).
[7] GIRSHICK R, DONAHUE J, DARRELL T, et al. Rich feature hierarchies for accurate object detection and se-mantic segmentation[C].//27th IEEE Conference on Com-puter Vision and Pattern Recognition (CVPR). Pisca-taway: IEEE Press,2014:580-587.
[8] GIRSHICK R. Fast R-CNN[C].//IEEE International Con-ference on Computer Vision(ICCV). Piscataway: IEEE Press,2015:1440-1448.
[9] REN S, HE K, GIRSHICK R ,et al. Faster R-CNN: To-wards Real-Time Object Detection with Region Proposal Networks[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2017,39,(6):1137-1149.
[10] REDMON J, DIVVALA S, GIRSHICK R, et al. You Only Look Once: Unified, Real-Time Object Detec-tion[C].//2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE Press,2016:779-788.
[11] REDMON J,FARHADI A. YOLOv3:An incremental improvement[EB/OL]:arXiv preprint:1804.02767,2018.
[12] JOCHER G, CHAURASIA A, QIU J. Ultralytics YOLO (Version 8.0.0)[EB/OL]. 2023:Available: https://github.com/ultralytics/ultralytics
[13] LIU W, ANGUELOV D, ERHAN D, et al. SSD: Single Shot MultiBox Detector[C].//14th European Conference on Computer Vision (ECCV). Cham:Springer.2016:21-37.
[14] CARION N, MASSA F, SYNNAEVE G, et al. End-to-End Object Detection with Transformers[C].//16th Euro-pean Conference on Computer Vision(ECCV). Cham:Springer.2020:213-229.
[15] LIN T Y,DOLLáR P,GIRSHICK R,et al. Feature pyramid networks for object detection[C]2017 IEEE Conference on Computer Vision and Pattern Recognition(CVPR). Piscataway:IEEE Press,2017:936-944.
[16] LIN T Y,GOYAL P,GIRSHICK R,et al. Focal loss for dense object detection[C]∥2017 IEEE Internation-al Conference on Computer Vision(ICCV).Piscataway:IEEE Press,2017:2999-3007
[17] DING J,XUE N,LONG Y,et al. Learning RoI transformer for oriented object detection in aerial images[C]∥2019 IEEE/CVF Conference on Computer Vsion and Pattern Recognition(CVPR). Piscataway: IEEE Press,2019:2844-2853.
[18] HAN J M, DING J, XUE N, et al. ReDet: A Rotation-equivariant Detector for Aerial Object Detec-tion[C].//IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE Press,2021:2785-2794.
[19] 于傲泽,魏维伟,王平,等. 基于分块复合注意力的无人机小目标检测算法[J].航空学报,2024,45(14):629148.YUA Z,WEI W W,WANG P,et al. Small target detection algorithm for UAV based on patch-wise co-attention[J].Acta Aeronautica et As-tronautica Sinica,2024,45(14):629148(in Chi-nese).
[20] 李子豪,王正平,贺云涛. 基于自适应协同注意力机制的航拍密集小目标检测算法[J].航空学报,2023,44(13):327944.LI Z H,WANG Z P,HE Y T. Aerial-photography dense small target detection algo-rithm based on adaptive cooperative attention mechanism[J].Acta Aeronautica et Astronautica Sinica,2023,44(13):327944(in Chinese).
[21] 刘延芳,佘佳宇,袁秋帆等.无人机遥感图像实时小目标检测方法[J].航空学报,2024,45,(14):53-72.LIU YF, SHE J Y,YUAN Q F,et al. Real-time small target de-tection networks for UAV remote sensing[J].Acta Aeronautica et Astronautica Sinica,2024,45(14):630119(in Chinese).
[22] XIAO Y, XU T F, XIN Y, et al. FBRT-YOLO: Faster and Better for Real-Time Aerial Image Detection[C].//39th AAAI Conference on Artificial Intelligence.2025:8673-8681.
[23] LI C L, YANG T J N, ZHU S J, et al. Density Map Guided Object Detection in Aerial Images[C].//IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE Press,2020:737-746.
[24] ZHANG Y, YE M, ZHU G Y, et al. FFCA-YOLO for Small Object Detection in Remote Sensing Imag-es[J].IEEE Transactions on Geoscience and Remote Sens-ing,2024,62:1-14.
[25] HU L, YUAN J W, CHENG B L, et al. CSFPR-RTDETR: Real-Time Small Object Detection Network for UAV Images Based on Cross-Spatial-Frequency Domain and Position Relation[J].IEEE Transactions on Geosci-ence and Remote Sensing,2025,63:3601828.
[26] TANG L F, YUAN J T, ZHANG H, et al. PIAFusion: A progressive infrared and visible image fusion network based on illumination aware[J].Information Fu-sion,2022,83-84,79-92.
[27] 常凯旋,黄建华,孙希延等.基于双模态图像融合的无人机光学小目标检测算法[J].激光与光电子学进展,2025,62,(4):269-283. CHANG K X,HUANG J H, SUN X Y,et al. Optical Small Target Detection Method by Drone Based on Dual-Modal Image Fusion[J].Laser & Optoelectronics Progress, 2025,62,(4): 269-283(in Chinese).
[28] FANG Q Y, HAN D P, WANG Z K. Cross-Modality Fusion Transformer for Multispectral Object Detec-tion[DB/OL].arXiv preprint.2111.00273,2021.
[29] ZENG Y Q, LIANG T F, JIN Y, et al. MMI-Det: Explor-ing Multi-Modal Integration for Visible and Infrared Ob-ject Detection[J].IEEE Transactions on Circuits and Sys-tems for Video Technology,2024,34,(11):11198-11213.
[30] YUAN M X, WEI X X.C2Former: Calibrated and Com-plementary Transformer for RGB-Infrared Object Detec-tion[J].IEEE Transactions on Geoscience and Remote Sensing, 2024, 62:1-12.
[31] GUO J J, GAO C Q, LIU F C, et al. DPDETR: Decou-pled Position Detection Transformer for Infrared-Visible Object Detection[DB/OL].arXiv pre-print.2408.06123.2025.
[32] DONG W H, ZHU H D, LIN S H , et al. Fusion-Mamba for Cross-Modality Object Detection[J].IEEE Transac-tions on Multimedia,2025,27,7392-7406.
[33] ZHOU M H, LI T Y, QIAO C F, et al. DMM: Disparity-Guided Multispectral Mamba for Oriented Object Detec-tion in Remote Sensing[J].IEEE Transactions on Geosci-ence and Remote Sensing,2025,63:3578309.
[34] YU C S, SHIN Y A. A Cross-Modality Feature Adaptive Interaction Approach for RGB-Infrared Object Detection in Aerial Imagery[J].IEEE Transactions on Geoscience and Remote Sensing,2026,64:3657379.
[35] ZHU P, SUN Y, WEN L, et al. Drone Based RGBT Ve-hicle Detection and Counting: A Challenge[DB/OL].arXiv preprint.2003.02437.2020.
[36] RAZAKARIVONY S, JURIE F. Vehicle detection in aerial imagery : A small target detection bench-mark[J].Journal of Visual Communication and Image Representation,2016,34:187-203.
文章导航

/