Information Fusion

Long-tailed ship recognition method based on aerial-space multimodal perception

  • Shihao WANG ,
  • Zhengwei XU ,
  • Long GAO ,
  • Congan XU ,
  • Yun LIN
Expand
  • 1.College of Information and Communication Engineering,Harbin Engineering University,Harbin 150001,China
    2.College of Aeronautical Science and Engineering,Henan Normal University,Xinxiang 453007,China
    3.Naval Aeronautical University,Yantai 264001,China

Received date: 2025-10-16

  Revised date: 2025-10-22

  Accepted date: 2025-11-04

  Online published: 2025-11-07

Supported by

National Natural Science Foundation of China(U23A20271);Special Funds for Basic Scientific Research Operations of Central Universities(3072025YY0801);China Postdoctoral Science Foundation(GZC20233554);Taishan Scholar Program(tsqn202312258)

Abstract

In the context of integrated aerial-space ocean monitoring and intelligent maritime management, ship target recognition based on multisource sensing data collected from Unmanned Aerial Vehicles (UAVs) and satellites plays a crucial role in navigation control, maritime law enforcement, and border surveillance. However, real-world ship recognition tasks face two major challenges. First, multimodal data fusion is difficult due to the heterogeneity and spatiotemporal misalignment between different modalities, such as optical images and electromagnetic radiation signals. Second, ship categories naturally exhibit a severe long-tailed distribution, where head classes dominate the sample population while tail classes remain scarce, significantly degrading overall recognition performance. To address these challenges, this paper proposes a long-tailed ship recognition method oriented toward aerial-space multimodal perception. The proposed method integrates a class-aware boundary optimization strategy and a category-based reweighting mechanism, effectively enhancing the discriminative capability of tail classes and improving the robustness of multimodal fusion. Experimental results demonstrate that the proposed method consistently outperforms existing approaches on representative long-tailed ship recognition tasks, showing strong practicality and generalization capability.

Cite this article

Shihao WANG , Zhengwei XU , Long GAO , Congan XU , Yun LIN . Long-tailed ship recognition method based on aerial-space multimodal perception[J]. ACTA AERONAUTICAET ASTRONAUTICA SINICA, 2026 , 47(S1) : 732925 -732925 . DOI: 10.7527/S1000-6893.2025.32925

References

[1] GUO L T, WANG Y, LIU Y C, et al. Ultralight convolutional neural network for automatic modulation classification in Internet of unmanned aerial vehicles[J]. IEEE Internet of Things Journal202411(11): 20831-20839.
[2] LIANG P P, LING C K, CHENG Y, et al. Multimodal learning without labeled multimodal data: Guarantees and applications[C]∥ International Conference on Learning Representations, 2024.
[3] 欧阳昱中,韩锐,刘驰.边缘侧领域自适应中长尾视觉识别技术研究[J].计算机工程202551(7): 171-179.
  OUYANG Y Z, HAN R, LIU C. Research on long-tail visual recognition technology with edge-side domain adaptation[J]. Computer Engineering202551(7): 171-179 (in Chinese).
[4] HE X Y, WANG Y, ZHAO S, et al. Co-attention fusion network for multimodal skin cancer diagnosis[J]. Pattern Recognition2023133: 108990.
[5] LU Y X, ZHAO W C, SUN N, et al. Enhancing multimodal knowledge graph representation learning through triple contrastive learning[C]∥ Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence. New York: ACM, 2024: 5963-5971.
[6] HUANG C, CAI W C, JIANG Q P, et al. Multimodal representation distribution learning for medical image segmentation[C]∥ Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence. New York: ACM, 2024: 4156-4164.
[7] ZHANG X C, DEMIRIS Y. Visible and infrared image fusion using deep learning[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence202345(8): 10535-10554.
[8] 郭浩, 李欣奕, 唐九阳, 等. 自适应特征融合的多模态实体对齐研究[J]. 自动化学报202450(4): 758-770.
  GUO H, LI X Y, TANG J Y, et al. Adaptive feature fusion for multi-modal entity alignment[J]. Acta Automatica Sinica202450(4): 758-770 (in Chinese).
[9] MA H Y, HE D X, WANG X B, et al. Multi-modal sarcasm detection based on dual generative processes[C]∥ International Joint Conference on Artificial Intelligence, 2024
[10] WANG J P, XU C A, ZHAO C H, et al. Multimodal object detection of UAV remote sensing based on joint representation optimization and specific information enhancement[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing202417: 12364-12373.
[11] WANG Q W, YIN C, SONG H C, et al. UTFNet: Uncertainty-guided trustworthy fusion network for RGB-thermal semantic segmentation[J]. IEEE Geoscience and Remote Sensing Letters202320: 7001205.
[12] 韩佳艺, 刘建伟, 陈德华, 等. 深度长尾学习研究综述[J]. 自动化学报202551(5): 985-1020.
  HAN J Y, LIU J W, CHEN D H, et al. Survey on deep long-tailed learning[J]. Acta Automatica Sinica202551(5): 985-1020 (in Chinese).
[13] ZHU J, WANG Z, CHEN J, et al. Balanced contrastive learning for long-tailed visual recognition[C]∥ Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022: 6908-6917.
[14] WANG Q W, QU X, JIN P C, et al. ODinMJ: A red, green, blue-thermal dataset for mountain jungle object detection[J]. IEEE Geoscience and Remote Sensing Magazine202513(3): 496-511.
[15] ZHANG X H, YOON J, BANSAL M, et al. Multimodal representation learning by alternating unimodal adaptation[C]∥ 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE Press, 2024: 27446-27456.
[16] 魏秀参, 许玉燕, 杨健. 网络监督数据下的细粒度图像识别综述[J]. 中国图象图形学报202227(7): 2057-2077.
  WEI X S, XU Y Y, YANG J. Review of webly-supervised fine-grained image recognition[J]. Journal of Image and Graphics202227(7): 2057-2077 (in Chinese).
[17] TU Y, LIN Y, HOU C B, et al. Complex-valued networks for automatic modulation classification[J]. IEEE Transactions on Vehicular Technology202069(9): 10085-10089.
[18] GUO L T, LIU C, LIU Y C, et al. Toward open-set specific emitter identification using auxiliary classifier generative adversarial network and OpenMax[J]. IEEE Transactions on Cognitive Communications and Networking202410(6): 2019-2028.
[19] ZHANG Y D, LATHAM P E, SAXE A. Understanding unimodal bias in multimodal deep linear networks[DB/OL]. arXiv preprint: 2312.00935, 2023.
[20] WANG S H, XU Z W, LIN Y. Multistage training and fusion method for imbalanced multimodal UAV remote sensing classification[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing202518: 18240-18250.
[21] LIN T Y, GOYAL P, GIRSHICK R, et al. Focal loss for dense object detection[C]∥ 2017 IEEE International Conference on Computer Vision (ICCV). Piscataway: IEEE Press, 2017: 2999-3007.
[22] CAO K D, WEI C, GAIDON A, et al. Learning imbalanced datasets with label-distribution-aware margin loss[C]∥ Neural Information Processing Systems, 2019.
Outlines

/