首页 >

跨模态注意力交互的航拍小目标智能检测算法-AI+空天科学

唐鼎新1,杨文青2,符强1,宣建林2   

  1. 1. 西北工业大学航空学院
    2. 西北工业大学
  • 收稿日期:2026-03-30 修回日期:2026-05-25 出版日期:2026-06-01 发布日期:2026-06-01
  • 通讯作者: 宣建林
  • 基金资助:
    715科技创新基金

Intelligent small object detection in aerial images via cross-modal attention interaction

  • Received:2026-03-30 Revised:2026-05-25 Online:2026-06-01 Published:2026-06-01

摘要: 在无人机航拍小目标检测任务中,如何在抑制模态间干扰的同时高效融合跨模态互补信息,是提升复杂环境下检测性能的关键问题。针对该挑战,本文提出一种双模态融合小目标检测算法PCMA-Net。首先,设计并行注意力与门控重标定模块,通过并行双路架构捕获细节特征,并利用动态门控掩码对增强特征执行二次特征精炼,抑制复杂背景伪影,提升模型在复杂背景下小目标判别能力。其次,构建跨模态注意力交互模块通过预建模响应分布聚焦显著区域,利用双向交叉注意力实现异质特征的像素级动态对齐,解决特征失配和掩没问题。最后,引入双分支注意力引导卷积,部分通道生成的动态权重引导主干路径进行特征增强,在较低计算开销下实现全局上下文信息建模,从而提升推理效率。实验结果表明,在DroneVehicle数据集上,所提方法相较于基线与主流融合方法,mAP@0.5分别提升1.9%和1.7%;在VEDAI数据集上较基线提升5.6%,并在多个类别上取得最优性能。同时,PCMA-Net实现了45.1 FPS的推理速度,满足无人机实时检测需求。

关键词: 无人机航拍图像, 小目标检测, 可见光-红外, 注意力机制, 轻量化设计

Abstract: In unmanned aerial vehicle (UAV) aerial photography tasks, the core challenge for achieving precision detection in complex environments lies in effectively suppressing inter-modal interference while efficiently fusing cross-modal complementary features. To address this challenge, this paper proposes PCMA-Net, a dual-modal fusion algorithm for small object detection in UAV imagery. First, a parallel attention and gated recalibration module is designed. By capturing detailed features through a parallel dual-path architecture and employing dynamic gated masking for secondary feature refinement, this module suppresses complex background artifacts and enhances the model's discriminative power for small objects in complex backgrounds. Second, a cross-modal attention interaction module is constructed. By pre-modeling the response distribution to focus on salient regions, it utilizes bidirectional cross-attention to achieve pixel-level dynamic alignment of heterogeneous features, thereby resolving feature mismatch and submergence issues. Finally, we propose dual-branch attention guided convolution, which utilizes channel splitting and task-aware weight generation to achieve global context modeling with minimal overhead, thereby improving inference efficiency on embedded platforms. Experimental results on the DroneVehicle dataset demonstrate that the proposed method outperforms the baseline and state-of-the-art fusion algorithms by 1.9% and 1.7% in mAP@0.5, respectively. Furthermore, on the VEDAI dataset, PCMA-Net exceeds the baseline by 5.6% in mAP@0.5 and achieves state-of-the-art performance across most specific categories. Notably, PCMA-Net achieves a detection speed of 45.1 FPS, fully meeting the real-time requirements for UAV-based deployment.

Key words: aerial imagery, small object detection, visible-infrared, attention mechanism, lightweight design

中图分类号: