基于空间关系驱动的遥感视觉定位方法-AI+空天科学

  • 刘崇基 ,
  • 张一鸣 ,
  • 刘瑜 ,
  • 姜智卓 ,
  • 毛永强 ,
  • 张霖平 ,
  • 何友
展开
  • 1. 清华大学深圳国际研究生院
    2. 清华大学电子工程系
    3. 南开大学
    4. 海军航空大学

收稿日期: 2026-04-02

  修回日期: 2026-07-20

  网络出版日期: 2026-07-22

基金资助

博士后基金;博士后创新人才支持;国家自然科学基金项目;国家自然科学基金项目;国家自然科学基金项目;国家自然科学基金项目

Spatial Relation-Driven Remote Sensing Visual Grounding Method

  • LIU Chong-Ji ,
  • ZHANG Yi-Ming ,
  • LIU Yu ,
  • JIANG Zhi-Zhuo ,
  • MAO Yong-Qiang ,
  • ZHANG Lin-Ping ,
  • HE You
Expand

Received date: 2026-04-02

  Revised date: 2026-07-20

  Online published: 2026-07-22

摘要

遥感视觉定位旨在通过自然语言查询实现影像中目标地物的精准定位,是遥感影像智能解译的关键技术。现有遥感视觉定位方法面临两类核心瓶颈:一是训练数据中空间关系标注粗糙且语义单一,导致模型难以从弱标注文本中习得可靠的空间推理能力;二是缺乏对复杂查询中语义主次关系的显式建模,在语义冗杂或参照物干扰显著的场景下,模型易将干扰项误判为定位主体。针对上述问题,本文提出基于空间关系驱动的遥感视觉定位方法SRMVG。首先,构建了包含精细空间关系标注的DIOR-RSVG-SR数据集,提供高质量的空间语义监督信号;其次,在跨模态推理层面提出基于双分支图神经网络的空间关系对齐框架,通过对文本特征进行四要素结构化解析实现图文对齐,并引入地理方位约束掩码以校正空间语义错配;最后,在目标精细定位方面提出语义优先级引导的筛选模块,通过自适应权重融合机制抑制参照物干扰引发的目标误判。实验结果表明,所提方法在DIOR-RSVG和OPT-RSVG两个公开数据集上均取得了优于现有方法的综合性能;在构建的DIOR-RSVG-SR数据集上,SRMVG的提升幅度显著优于其他对比方法,验证了本文方法在利用精细空间监督信号方面的优越性。

本文引用格式

刘崇基 , 张一鸣 , 刘瑜 , 姜智卓 , 毛永强 , 张霖平 , 何友 . 基于空间关系驱动的遥感视觉定位方法-AI+空天科学[J]. 航空学报, 0 : 1 -0 . DOI: 10.7527/S1000-6893.2026.33673

Abstract

Remote sensing visual grounding aims to precisely locate target ground objects in images via natural language queries, serving as a critical technology for intelligent interpretation of remote sensing images. Existing remote sensing visual grounding methods suffer from two core bottlenecks. First, coarse spatial relation annotations and monotonous semantics in training data hinder models from learning robust spatial reasoning capabilities from weakly annotated texts. Second, explicit modeling of primary and secondary semantic relationships within complex queries is absent; in scenarios with redundant semantics or severe reference object interference, the model tends to misclassify distractors as localization targets. To address the above issues, this paper proposes a spatial relation-driven remote sensing visual grounding method named SRMVG. First, a DIOR-RSVG-SR dataset with elaborate spatial relation annotations is constructed to provide high-quality spatial semantic supervision signals. Second, at the cross-modal reasoning stage, a spatial relation alignment framework based on two-branch graph neural networks is proposed. It realizes text-image alignment through structured parsing of four core elements from text features, and introduces a geographic orientation constraint mask to rectify mismatched spatial semantics. Finally, a semantic priority-guided filtering module is designed for fine-grained target localization, which suppresses target misjudgments caused by reference object interference via an adaptive weight fusion mechanism. Experimental results demonstrate that the proposed method achieves overall performance superior to state-of-the-art approaches on two public datasets, DIOR-RSVG and OPT-RSVG. On the self-built DIOR-RSVG-SR dataset, SRMVG achieves remarkably greater performance gains compared with other competitive methods, which verifies the superiority of the proposed method in leveraging elaborate spatial supervision signals.

参考文献

[1]ZHAN Y, XIONG Z, YUAN Y.RSVG: Exploring data and models for visual grounding on remote sensing data[J]. IEEE Transactions on Geoscience and Remote Sensing, 2023, 61: 1-13.[J].IEEE Transactions on Geoscience and Remote Sensing, 2023, 61(无):1-13 [2]SUN Y, FENG S, LI X, et al.Visual grounding in remote sensing images[C]//Proceedings of the ACM International Conference on Multimedia. New York: ACM, 2022: 404-412..Proceedings of the ACM International Conference on Multimedia, 2022, 无(无):404-412 [3]LI T, ZHANG Y, WANG C, et al.TACMT: Text-aware cross-modal transformer for visual grounding on high-resolution SAR images[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2025, 222: 152-166.[J].ISPRS Journal of Photogrammetry and Remote Sensing, 2025, 222(无):152-166 [4]MAO J, HUANG J, TOSHEV A, et al.Generation and comprehension of unambiguous object descriptions[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE, 2016: 11-20..Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, 无(无):11-20 [5]YANG Z, GONG B, WANG L, et al.A joint speaker-listener-reinforcer model for referring expression comprehension[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE, 2019: 12404-12413..Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, 无(无):12404-12413 [6]DENG J, YANG Z, CHEN T, et al.TransVG: End-to-end visual grounding with transformers[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Piscataway: IEEE, 2021: 1769-1779..Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, 无(无):1769-1779 [7]YANG L, XU Y, YUAN C, et al.Improving visual grounding with visual-linguistic verification and iterative reasoning[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE, 2022: 9499-9508..Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, 无(无):9499-9508 [8]WANG F, WU C, WU J, et al.Multistage synergistic aggregation network for remote sensing visual grounding[J]. IEEE Geoscience and Remote Sensing Letters, 2024, 21: 1-5.[J].IEEE Geoscience and Remote Sensing Letters, 2024, 21(无):1-5 [9]DING Y, WANG D, LI K, et al.Visual grounding of remote sensing images with multi-dimensional semantic-guidance[J]. Pattern Recognition Letters, 2025, 189: 85-91.[J].Pattern Recognition Letters, 2025, 189(无):85-91 [10]MA Q, PAN J, BAI C.Direction-oriented visual-semantic embedding model for remote sensing image-text retrieval[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 4704014.[J].IEEE Transactions on Geoscience and Remote Sensing, 2024, 62(无):1-14 [11]WANG M, GUO J, SONG B, et al.Graph-based hierarchical semantic consistency network for remote sensing image-text retrieval[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2025, 18: 15334-15346.[J].IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2025, 18(无):15334-15346 [12]LI J, ZHANG Y, WANG C, et al.Multi-scale cross-modal attention fusion for remote sensing visual grounding[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2024, 17: 3215-3228.[J].IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2024, 17(无):3215-3228 [13]YUAN Z, ZHANG W, RONG X, et al.A lightweight multi-scale crossmodal text-image retrieval method in remote sensing[J]. IEEE Transactions on Geoscience and Remote Sensing, 2021, 60: 1-19.[J].IEEE Transactions on Geoscience and Remote Sensing, 2021, 60(无):1-19 [14]CHENG Q, ZHOU Y, FU P, et al.A deep semantic alignment network for the cross-modal image-text retrieval in remote sensing[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2021, 14: 4284-4297.[J].IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2021, 14(无):4284-4297 [15]YU H, YAO F, LU W, et al.Text-image matching for cross-modal remote sensing image retrieval via graph neural network[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2022, 16: 812-824.[J].IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2022, 16(无):812-824 [16]FU K, ZHANG X, LIU G, et al.Scattering-keypoint-guided network for oriented ship detection in high-resolution and large-scale SAR images[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2021, 14: 11162-11178.[J].IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2021, 14(无):11162-11178 [17]LU X, WANG B, ZHENG X, et al.Exploring models and data for remote sensing image caption generation[J].IEEE Transactions on Geoscience and Remote Sensing, 2017, 56(4):2183-2195 [18]LI Y, MAO H, GIRSHICK R, et al.Exploring plain vision transformer backbones for object detection[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE, 2022: 2804-2814..Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, 无(无):2804-2814 [19]LIU Y, OTT M, GOYAL N, et al.RoBERTa: A robustly optimized BERT pretraining approach[EB/OL]. (2019-07-26)[2026-03-30]. https://arxiv.org/abs/1907.11692.[J].arXiv preprint arXiv:1907.11692, 2019, 无(无):无-无 [20]REN S, HE K, GIRSHICK R, et al.Faster R-CNN: Towards real-time object detection with region proposal networks[J].IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 39(6):1137-1149 [21]YANG Z, GONG B, WANG L, et al.A fast and accurate one-stage approach to visual grounding[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Piscataway: IEEE, 2019: 4683-4693..Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, 无(无):4683-4693 [22]YANG Z, CHEN T, WANG L, et al.Improving one-stage visual grounding by recursive sub-query construction[C]//Proceedings of the European Conference on Computer Vision. Cham: Springer, 2020: 387-404..Proceedings of the European Conference on Computer Vision (ECCV), 2020, 无(无):387-404 [23]HUANG B, LIAN D, LUO W, et al.Look before you leap: learning landmark features for one-stage visual grounding[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE, 2021: 16888-16897..Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, 无(无):16888-16897
Options
文章导航

/