ACTA AERONAUTICAET ASTRONAUTICA SINICA >
Long-tailed ship recognition method based on aerial-space multimodal perception
Received date: 2025-10-16
Revised date: 2025-10-22
Accepted date: 2025-11-04
Online published: 2025-11-07
Supported by
National Natural Science Foundation of China(U23A20271);Special Funds for Basic Scientific Research Operations of Central Universities(3072025YY0801);China Postdoctoral Science Foundation(GZC20233554);Taishan Scholar Program(tsqn202312258)
In the context of integrated aerial-space ocean monitoring and intelligent maritime management, ship target recognition based on multisource sensing data collected from Unmanned Aerial Vehicles (UAVs) and satellites plays a crucial role in navigation control, maritime law enforcement, and border surveillance. However, real-world ship recognition tasks face two major challenges. First, multimodal data fusion is difficult due to the heterogeneity and spatiotemporal misalignment between different modalities, such as optical images and electromagnetic radiation signals. Second, ship categories naturally exhibit a severe long-tailed distribution, where head classes dominate the sample population while tail classes remain scarce, significantly degrading overall recognition performance. To address these challenges, this paper proposes a long-tailed ship recognition method oriented toward aerial-space multimodal perception. The proposed method integrates a class-aware boundary optimization strategy and a category-based reweighting mechanism, effectively enhancing the discriminative capability of tail classes and improving the robustness of multimodal fusion. Experimental results demonstrate that the proposed method consistently outperforms existing approaches on representative long-tailed ship recognition tasks, showing strong practicality and generalization capability.
Shihao WANG , Zhengwei XU , Long GAO , Congan XU , Yun LIN . Long-tailed ship recognition method based on aerial-space multimodal perception[J]. ACTA AERONAUTICAET ASTRONAUTICA SINICA, 2026 , 47(S1) : 732925 -732925 . DOI: 10.7527/S1000-6893.2025.32925
| [1] | GUO L T, WANG Y, LIU Y C, et al. Ultralight convolutional neural network for automatic modulation classification in Internet of unmanned aerial vehicles[J]. IEEE Internet of Things Journal, 2024, 11(11): 20831-20839. |
| [2] | LIANG P P, LING C K, CHENG Y, et al. Multimodal learning without labeled multimodal data: Guarantees and applications[C]∥ International Conference on Learning Representations, 2024. |
| [3] | 欧阳昱中,韩锐,刘驰.边缘侧领域自适应中长尾视觉识别技术研究[J].计算机工程, 2025, 51(7): 171-179. |
| OUYANG Y Z, HAN R, LIU C. Research on long-tail visual recognition technology with edge-side domain adaptation[J]. Computer Engineering, 2025, 51(7): 171-179 (in Chinese). | |
| [4] | HE X Y, WANG Y, ZHAO S, et al. Co-attention fusion network for multimodal skin cancer diagnosis[J]. Pattern Recognition, 2023, 133: 108990. |
| [5] | LU Y X, ZHAO W C, SUN N, et al. Enhancing multimodal knowledge graph representation learning through triple contrastive learning[C]∥ Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence. New York: ACM, 2024: 5963-5971. |
| [6] | HUANG C, CAI W C, JIANG Q P, et al. Multimodal representation distribution learning for medical image segmentation[C]∥ Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence. New York: ACM, 2024: 4156-4164. |
| [7] | ZHANG X C, DEMIRIS Y. Visible and infrared image fusion using deep learning[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023, 45(8): 10535-10554. |
| [8] | 郭浩, 李欣奕, 唐九阳, 等. 自适应特征融合的多模态实体对齐研究[J]. 自动化学报, 2024, 50(4): 758-770. |
| GUO H, LI X Y, TANG J Y, et al. Adaptive feature fusion for multi-modal entity alignment[J]. Acta Automatica Sinica, 2024, 50(4): 758-770 (in Chinese). | |
| [9] | MA H Y, HE D X, WANG X B, et al. Multi-modal sarcasm detection based on dual generative processes[C]∥ International Joint Conference on Artificial Intelligence, 2024 |
| [10] | WANG J P, XU C A, ZHAO C H, et al. Multimodal object detection of UAV remote sensing based on joint representation optimization and specific information enhancement[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2024, 17: 12364-12373. |
| [11] | WANG Q W, YIN C, SONG H C, et al. UTFNet: Uncertainty-guided trustworthy fusion network for RGB-thermal semantic segmentation[J]. IEEE Geoscience and Remote Sensing Letters, 2023, 20: 7001205. |
| [12] | 韩佳艺, 刘建伟, 陈德华, 等. 深度长尾学习研究综述[J]. 自动化学报, 2025, 51(5): 985-1020. |
| HAN J Y, LIU J W, CHEN D H, et al. Survey on deep long-tailed learning[J]. Acta Automatica Sinica, 2025, 51(5): 985-1020 (in Chinese). | |
| [13] | ZHU J, WANG Z, CHEN J, et al. Balanced contrastive learning for long-tailed visual recognition[C]∥ Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022: 6908-6917. |
| [14] | WANG Q W, QU X, JIN P C, et al. ODinMJ: A red, green, blue-thermal dataset for mountain jungle object detection[J]. IEEE Geoscience and Remote Sensing Magazine, 2025, 13(3): 496-511. |
| [15] | ZHANG X H, YOON J, BANSAL M, et al. Multimodal representation learning by alternating unimodal adaptation[C]∥ 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway: IEEE Press, 2024: 27446-27456. |
| [16] | 魏秀参, 许玉燕, 杨健. 网络监督数据下的细粒度图像识别综述[J]. 中国图象图形学报, 2022, 27(7): 2057-2077. |
| WEI X S, XU Y Y, YANG J. Review of webly-supervised fine-grained image recognition[J]. Journal of Image and Graphics, 2022, 27(7): 2057-2077 (in Chinese). | |
| [17] | TU Y, LIN Y, HOU C B, et al. Complex-valued networks for automatic modulation classification[J]. IEEE Transactions on Vehicular Technology, 2020, 69(9): 10085-10089. |
| [18] | GUO L T, LIU C, LIU Y C, et al. Toward open-set specific emitter identification using auxiliary classifier generative adversarial network and OpenMax[J]. IEEE Transactions on Cognitive Communications and Networking, 2024, 10(6): 2019-2028. |
| [19] | ZHANG Y D, LATHAM P E, SAXE A. Understanding unimodal bias in multimodal deep linear networks[DB/OL]. arXiv preprint: 2312.00935, 2023. |
| [20] | WANG S H, XU Z W, LIN Y. Multistage training and fusion method for imbalanced multimodal UAV remote sensing classification[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2025, 18: 18240-18250. |
| [21] | LIN T Y, GOYAL P, GIRSHICK R, et al. Focal loss for dense object detection[C]∥ 2017 IEEE International Conference on Computer Vision (ICCV). Piscataway: IEEE Press, 2017: 2999-3007. |
| [22] | CAO K D, WEI C, GAIDON A, et al. Learning imbalanced datasets with label-distribution-aware margin loss[C]∥ Neural Information Processing Systems, 2019. |
/
| 〈 |
|
〉 |