电子电气工程与控制

基于图卷积网络的红外目标检测算法

  • 仝照亚 ,
  • 刘刚 ,
  • 霍元智 ,
  • 樊肖亮 ,
  • 吕书贤
展开
  • 河南科技大学 信息工程学院,洛阳 471023
.E-mail: gliu@haust.edu.cn

收稿日期: 2025-10-10

  修回日期: 2025-11-25

  录用日期: 2025-12-19

  网络出版日期: 2025-12-25

基金资助

国家留学基金([2022]20)

Infrared target detection algorithm based on graph convolutional network

  • Zhaoya TONG ,
  • Gang LIU ,
  • Yuanzhi HUO ,
  • Xiaoliang FAN ,
  • Shuxian LYU
Expand
  • School of Information Engineering,Henan University of Science and Technology,Luoyang 471023,China
E-mail: gliu@haust.edu.cn

Received date: 2025-10-10

  Revised date: 2025-11-25

  Accepted date: 2025-12-19

  Online published: 2025-12-25

Supported by

National Scholarship Foundation of China([2022]20)

摘要

红外目标因背景干扰强、纹理信息弱及结构特征模糊等特性导致其检测精度受限,且现有基于深度学习的方法多依赖单目标特征,忽略了目标间的关联信息。为此,提出一种结合类间关联和语义相似性的图卷积网络红外目标检测算法,其核心在于构建动态融合的图结构。首先,利用目标在图像中的共现概率构建静态的共现邻接矩阵,以捕获目标之间的上下文关系。同时,基于预训练的词向量构建语义邻接矩阵,并在训练过程中对其进行更新,使其能够动态适应目标标签的语义特性。随后,将2类邻接矩阵进行融合,并以类别词向量作为节点特征输入图卷积网络,从而实现对类别关系的高阶建模。最终,所获得的关系特征通过注意力机制与YOLO11主干网络提取的特征进行融合,用于后续检测头的分类分支和定位分支。在自制红外飞机数据集上得到的实验结果表明,与YOLO11相比,在计算量和参数量相当的条件下,所提算法将mAP50和mAP50:95分别提升了1.08%和0.86%,与其他最新的SOTA算法相比,所提算法的检测精度同样达到了最优。理论分析和实验结果共同验证了所提算法在复杂环境下执行红外目标检测任务的有效性。

本文引用格式

仝照亚 , 刘刚 , 霍元智 , 樊肖亮 , 吕书贤 . 基于图卷积网络的红外目标检测算法[J]. 航空学报, 2026 , 47(12) : 332886 -332886 . DOI: 10.7527/S1000-6893.2025.32886

Abstract

Infrared target detection is limited by strong background interference, weak texture information, and blurred structural features. Existing deep learning-based methods mostly rely on single-target features, neglecting inter-target correlation information. To address this issue, this paper proposes a graph convolutional network-based infrared target detection algorithm that integrates inter-class correlation and semantic similarity, with the core being the construction of a dynamically fused graph structure. Firstly, a static co-occurrence adjacency matrix is constructed using the co-occurrence probability of targets in the image to capture the contextual relationship between targets. Simultaneously, a semantic adjacency matrix is constructed based on pre-trained word vectors and updated during training to dynamically adapt to the semantic characteristics of target labels. Subsequently, the two adjacency matrices are fused, and the category word vectors are input into the graph convolutional network as node features, achieving high-level modeling of category relationship. Finally, the obtained relational features are fused with features extracted by the YOLO11 backbone network through an attention mechanism for subsequent classification and localization branches of the detection head. Experimental results on a self-made infrared aircraft dataset demonstrate that, compared with YOLO11, the proposed algorithm improves mAP50 and mAP50:95 by 1.08% and 0.86%, respectively, with comparable computational complexity and parameter count. Compared with other state-of-the-art algorithms, the proposed algorithm also achieves optimal detection accuracy. Theoretical analysis and experimental results demonstrate the effectiveness of the proposed algorithm in infrared target detection task under complex environments.

参考文献

[1] 于俊庭, 李少毅, 张平, 等. 光电成像末制导智能化技术研究与展望[J]. 红外与激光工程202352(5): 20220725.
  YU J T, LI S Y, ZHANG P, et al. Research and prospect of intelligent technology of optoelectronic imaging terminal guidance[J]. Infrared and Laser Engineering202352(5): 20220725 (in Chinese).
[2] 李少毅, 卫孟杰, 杨俊彦, 等. 红外多波段成像末制导技术研究现状与展望[J]. 航空学报202445(20): 630427.
  LI S Y, WEI M J, YANG J Y, et al. Research status and prospects of infrared multi-band imaging terminal guidance technology[J]. Acta Aeronautica et Astronautica Sinica202445(20): 630427 (in Chinese).
[3] 张阳婷, 黄德启, 王东伟, 等. 基于深度学习的目标检测算法研究与应用综述[J]. 计算机工程与应用202359(18): 1-13.
  ZHANG Y T, HUANG D Q, WANG D W, et al. Review on research and application of deep learning-based target detection algorithms[J]. Computer Engineering and Applications202359(18): 1-13 (in Chinese).
[4] 王强, 吴乐天, 王勇, 等. 基于关键点检测的红外弱小目标检测[J]. 航空学报202344(10): 328173.
  WANG Q, WU L T, WANG Y, et al. An infrared small target detection method based on key point[J]. Acta Aeronautica et Astronautica Sinica202344(10): 328173 (in Chinese).
[5] ZHOU J J, ZHANG B H, YUAN X L, et al. YOLO-CIR: The network based on YOLO and ConvNeXt for infrared object detection[J]. Infrared Physics Technology2023131: 104703.
[6] 王健宗, 孔令炜, 黄章成, 等. 图神经网络综述[J]. 计算机工程202147(4): 1-12.
  WANG J Z, KONG L W, HUANG Z C, et al. Survey of graph neural network[J]. Computer Engineering202147(4): 1-12 (in Chinese).
[7] JIA G M, CHENG Y, CHEN T. IRGraphSeg: Infrared small target detection based on hierarchical GNN[J]. IEEE Geoscience and Remote Sensing Letters202421: 6005505.
[8] YANG X, LI S Y, SUN S J, et al. Anti-occlusion infrared aerial target recognition with multisemantic graph skeleton model[J]. IEEE Transactions on Geoscience and Remote Sensing202260: 5629813.
[9] LIN J, LI S Y, YANG X, et al. CS-ViG-UNet: Infrared small and dim target detection based on cycle shift vision graph convolution network[J]. Expert Systems with Applications2024254: 124385.
[10] YANG R M, ZHANG Y D, GAO G S, et al. Global information aware network with global interaction graph attention for infrared small target detection[J]. IET Image Processing202418(12): 3650-3666.
[11] YE J, HE J J, PENG X J, et al. Attention-driven dynamic graph convolutional network for multi-label image recognition[C]∥ Computer Vision-ECCV 2020. Cham: Springer, 2020: 649-665.
[12] CHEN S J, LI Z X, HUANG F C, et al. Object detection using dual graph network[C]∥ 2020 25th International Conference on Pattern Recognition (ICPR). Piscataway: IEEE Press, 2021: 3280-3287.
[13] XU Z M, WEI L L, LANG C Y, et al. SSR-Net: A spatial structural relation network for vehicle re-identification[J]. ACM Transactions on Multimedia Computing, Communications, and Applications202319(6): 1-22.
[14] LI C, LIU F H, TIAN Z Q, et al. DAGCN: Dynamic and adaptive graph convolutional network for salient object detection[J]. IEEE Transactions on Neural Networks and Learning Systems202435(6): 7612-7626.
[15] WANG P C, TAO L P, TANG M W, et al. Incorporating syntax and semantics with dual graph neural networks for aspect-level sentiment analysis[J]. Engineering Applications of Artificial Intelligence2024133: 108101.
[16] SHI J Y, ZHONG J Q, CAO W M. Multi-semantics aggregation network based on the dynamic-attention mechanism for 3D human motion prediction[J]. IEEE Transactions on Multimedia202426: 5194-5206.
[17] WANG X Y, ZHANG L J, JIANG Y T, et al. Optimizing military target recognition in urban battlefields: An intelligent framework based on graph neural networks and YOLO[J]. Signal, Image and Video Processing202419(1): 6.
[18] CHEN A, ZHOU Y. An attention enhanced graph convolutional network for semantic segmentation[C]∥Pattern Recognition and Computer Vision. Cham: Springer, 2020: 734-745.
[19] LIU Y S, CHEN W Y, QU H, et al. Weakly supervised image classification and pointwise localization with graph convolutional networks[J]. Pattern Recognition2021109: 107596.
[20] XIE K Z, WEI Z Q, HUANG L, et al. Graph convolutional networks with attention for multi-label weather recognition[J]. Neural Computing and Applications202133(17): 11107-11123.
[21] ZHANG D, LI W H, QIU T B, et al. Co-occurrence graph convolutional networks with approximate entailment for knowledge graph embedding[J]. Applied Soft Computing2025170: 112666.
[22] ZHONG F J, SHEN W X, YU H, et al. Dehazing reasoning YOLO: Prior knowledge-guided network for object detection in foggy weather[J]. Pattern Recognition2024156: 110756.
[23] LIU Y K, NI K S, ZHANG Y H, et al. Semantic interleaving global channel attention for multilabel remote sensing image classification[J]. International Journal of Remote Sensing202445(2): 393-419.
[24] CHATTOPADHAY A, SARKAR A, HOWLADER P, et al. Grad-CAM++: Generalized gradient-based visual explanations for deep convolutional networks[C]∥ 2018 IEEE Winter Conference on Applications of Computer Vision (WACV). Piscataway: IEEE Press, 2018: 839-847.
[25] ZHANG S L, WANG X J, WANG J Q, et al. Dense distinct query for end-to-end object detection[C]∥ 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE Press, 2023: 7329-7338.
[26] ZHU X Z, SU W J, LU L W, et al. Deformable DETR: Deformable transformers for end-to-end object detection[DB/OL]. arXiv preprint2010.04159, 2020.
[27] ZHANG H K, CHANG H, MA B P, et al. Dynamic R-CNN: Towards high quality object detection via dynamic training[M]∥ Computer Vision-ECCV 2020. Cham: Springer International Publishing, 2020: 260-275.
[28] QIU H Q, LI H L, WU Q B, et al. CrossDet: Growing crossline representation for object detection[J]. IEEE Transactions on Circuits and Systems for Video Technology202333(3): 1093-1108.
[29] YAO R L, RONG Y, HUANG Q Q, et al. CTOD: Cross-attentive task-alignment for one-stage object detection[J]. IEEE Transactions on Circuits and Systems for Video Technology202434(11): 11507-11520.
[30] CHEN Z H, YANG C, CHANG J H, et al. DDOD: Dive deeper into the disentanglement of object detector[J]. IEEE Transactions on Multimedia202426: 284-298.
[31] LI Z H, DONG Y S. Refined feature enhancement network for object detection[J]. Complex Intelligent Systems202511(1): 1-15.
[32] TIAN Y J, YE Q X, DAVID D. YOLOv12: Attention-centric real-time object detectors[DB/OL]. arXiv preprint: 2502.12524, 2025.
[33] MUNIR M, AVERY W, RAHMAN M M, et al. GreedyViG: Dynamic axial graph construction for efficient vision GNNs[C]∥ 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE Press, 2024: 6118-6127.
[34] MUNIR M, AVERY W, MARCULESCU R. MobileViG: Graph-based sparse attention for mobile vision applications[C]∥ 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). Piscataway: IEEE Press, 2023: 2211-2219.
文章导航

/