无人机可见光-红外目标检测能够利用可见光图像的纹理细节与红外图像的热辐射信息,提升复杂环境下的目标感知能力。然而,现有方法多侧重跨模态特征交互与特征融合,忽略了融合前单模态特征中热噪声、伪热点、纹理冗余和背景边缘等高频干扰,导致高频噪声在跨模态传播中被进一步放大,降低检测精度。针对上述问题,提出了一种基于频率净化与交互融合的无人机可见光-红外目标检测方法,通过“先净化、后融合”的策略,实现单模态高频干扰抑制与跨模态互补信息增强。具体而言,设计了一个低频聚合频率净化(Low-frequency Aggregation Frequency Purification, LAFP)模块,利用跨模态低频结构先验引导各模态特征中的高频细节特征筛选与净化,在保留目标边缘与局部判别特征的同时抑制无效高频响应。此外,引入一个双向交互引导融合(Bidirectional Interactive Guidance Fusion, BIGF)模块,该模块利用双向交叉注意力挖掘净化后可见光与红外特征之间的互补关系,并结合动态门控机制实现跨模态特征自适应融合。在DroneVehicle、VEDAI和LLVIP三个可见光-红外目标检测数据集上的实验结果表明,本文方法在复杂航拍车辆检测与低照行人检测场景中均取得了良好的性能。相较现有先进方法,本文方法在DroneVehicle与VEDAI数据集上的平均精度均值(mean Average Precision,mAP)分别达到82.4%和80.5%,分别较次优方法提升3.0%和2.7%。在LLVIP数据集上取得96.1%的平均精度,虽然相较于最优和次优结果分别降低了0.7%和0.4%,但结果也具有较好的竞争力。实验结果表明,所提算法在无人机目标检测场景中具有较高的有效性和鲁棒性。
Visible-infrared object detection for unmanned aerial vehicles can exploit the texture details of visible images and the thermal radiation information of infrared images, thereby improving target perception capability in complex environments. However, existing methods mainly focus on cross-modal feature interaction and feature fusion, while ignoring high-frequency interference in single-modal features before fusion, such as thermal noise, pseudo-hotspots, texture redundancy, and background edges. As a result, high-frequency noise is further amplified during cross-modal propagation, reducing detection accuracy. To address these problems, this paper proposes a UAV visible-infrared object detection method based on frequency purification and interactive fusion. Through a “purification-before-fusion” strategy, the proposed method suppresses single-modal high-frequency interference and enhances cross-modal complementary information. Specifically, a Low-frequency Aggregation Frequency Purification (LAFP) module is designed to use cross-modal low-frequency structural priors to guide the selection and purification of high-frequency detail features in each modality, suppressing invalid high-frequency responses while preserving target edges and local discriminative features. In addition, a Bidirectional Interactive Guidance Fusion (BIGF) module is introduced. This module employs bidirectional cross-attention to mine the complementary relationships between purified visible and infrared features, and combines a dynamic gating mechanism to achieve adaptive cross-modal feature fusion. Experimental results on three visible-infrared object detection datasets, namely DroneVehicle, VEDAI, and LLVIP, show that the proposed method achieves good performance in both complex aerial vehicle detection and low-light pedestrian detection scenarios. Compared with existing state-of-the-art methods, the proposed method achieves mean Average Precision (mAP) values of 82.4% and 80.5% on the DroneVehicle and VEDAI datasets, respectively, outperforming the second-best methods by 3.0% and 2.7%. On the LLVIP dataset, it obtains an average precision of 96.1%, which is 0.7% and 0.4% lower than the best and second-best results, respectively, but still demonstrates strong competitiveness. The experimental results demonstrate that the proposed algorithm has high effectiveness and robustness in UAV object detection scenarios.