Remote sensing visual grounding aims to precisely locate target ground objects in images via natural language queries, serving as a critical technology for intelligent interpretation of remote sensing images. Existing remote sensing visual grounding methods suffer from two core bottlenecks. First, coarse spatial relation annotations and monotonous semantics in training data hinder models from learning robust spatial reasoning capabilities from weakly annotated texts. Second, explicit modeling of primary and secondary semantic relationships within complex queries is absent; in scenarios with redundant semantics or severe reference object interference, the model tends to misclassify distractors as localization targets. To address the above issues, this paper proposes a spatial relation-driven remote sensing visual grounding method named SRMVG. First, a DIOR-RSVG-SR dataset with elaborate spatial relation annotations is constructed to provide high-quality spatial semantic supervision signals. Second, at the cross-modal reasoning stage, a spatial relation alignment framework based on two-branch graph neural networks is proposed. It realizes text-image alignment through structured parsing of four core elements from text features, and introduces a geographic orientation constraint mask to rectify mismatched spatial semantics. Finally, a semantic priority-guided filtering module is designed for fine-grained target localization, which suppresses target misjudgments caused by reference object interference via an adaptive weight fusion mechanism. Experimental results demonstrate that the proposed method achieves overall performance superior to state-of-the-art approaches on two public datasets, DIOR-RSVG and OPT-RSVG. On the self-built DIOR-RSVG-SR dataset, SRMVG achieves remarkably greater performance gains compared with other competitive methods, which verifies the superiority of the proposed method in leveraging elaborate spatial supervision signals.
[1]ZHAN Y, XIONG Z, YUAN Y.RSVG: Exploring data and models for visual grounding on remote sensing data[J]. IEEE Transactions on Geoscience and Remote Sensing, 2023, 61: 1-13.[J].IEEE Transactions on Geoscience and Remote Sensing, 2023, 61(无):1-13
[2]SUN Y, FENG S, LI X, et al.Visual grounding in remote sensing images[C]//Proceedings of the ACM International Conference on Multimedia. New York: ACM, 2022: 404-412..Proceedings of the ACM International Conference on Multimedia, 2022, 无(无):404-412
[3]LI T, ZHANG Y, WANG C, et al.TACMT: Text-aware cross-modal transformer for visual grounding on high-resolution SAR images[J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2025, 222: 152-166.[J].ISPRS Journal of Photogrammetry and Remote Sensing, 2025, 222(无):152-166
[4]MAO J, HUANG J, TOSHEV A, et al.Generation and comprehension of unambiguous object descriptions[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE, 2016: 11-20..Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, 无(无):11-20
[5]YANG Z, GONG B, WANG L, et al.A joint speaker-listener-reinforcer model for referring expression comprehension[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE, 2019: 12404-12413..Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, 无(无):12404-12413
[6]DENG J, YANG Z, CHEN T, et al.TransVG: End-to-end visual grounding with transformers[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Piscataway: IEEE, 2021: 1769-1779..Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, 无(无):1769-1779
[7]YANG L, XU Y, YUAN C, et al.Improving visual grounding with visual-linguistic verification and iterative reasoning[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE, 2022: 9499-9508..Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, 无(无):9499-9508
[8]WANG F, WU C, WU J, et al.Multistage synergistic aggregation network for remote sensing visual grounding[J]. IEEE Geoscience and Remote Sensing Letters, 2024, 21: 1-5.[J].IEEE Geoscience and Remote Sensing Letters, 2024, 21(无):1-5
[9]DING Y, WANG D, LI K, et al.Visual grounding of remote sensing images with multi-dimensional semantic-guidance[J]. Pattern Recognition Letters, 2025, 189: 85-91.[J].Pattern Recognition Letters, 2025, 189(无):85-91
[10]MA Q, PAN J, BAI C.Direction-oriented visual-semantic embedding model for remote sensing image-text retrieval[J]. IEEE Transactions on Geoscience and Remote Sensing, 2024, 62: 4704014.[J].IEEE Transactions on Geoscience and Remote Sensing, 2024, 62(无):1-14
[11]WANG M, GUO J, SONG B, et al.Graph-based hierarchical semantic consistency network for remote sensing image-text retrieval[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2025, 18: 15334-15346.[J].IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2025, 18(无):15334-15346
[12]LI J, ZHANG Y, WANG C, et al.Multi-scale cross-modal attention fusion for remote sensing visual grounding[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2024, 17: 3215-3228.[J].IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2024, 17(无):3215-3228
[13]YUAN Z, ZHANG W, RONG X, et al.A lightweight multi-scale crossmodal text-image retrieval method in remote sensing[J]. IEEE Transactions on Geoscience and Remote Sensing, 2021, 60: 1-19.[J].IEEE Transactions on Geoscience and Remote Sensing, 2021, 60(无):1-19
[14]CHENG Q, ZHOU Y, FU P, et al.A deep semantic alignment network for the cross-modal image-text retrieval in remote sensing[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2021, 14: 4284-4297.[J].IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2021, 14(无):4284-4297
[15]YU H, YAO F, LU W, et al.Text-image matching for cross-modal remote sensing image retrieval via graph neural network[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2022, 16: 812-824.[J].IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2022, 16(无):812-824
[16]FU K, ZHANG X, LIU G, et al.Scattering-keypoint-guided network for oriented ship detection in high-resolution and large-scale SAR images[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2021, 14: 11162-11178.[J].IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2021, 14(无):11162-11178
[17]LU X, WANG B, ZHENG X, et al.Exploring models and data for remote sensing image caption generation[J].IEEE Transactions on Geoscience and Remote Sensing, 2017, 56(4):2183-2195
[18]LI Y, MAO H, GIRSHICK R, et al.Exploring plain vision transformer backbones for object detection[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE, 2022: 2804-2814..Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, 无(无):2804-2814
[19]LIU Y, OTT M, GOYAL N, et al.RoBERTa: A robustly optimized BERT pretraining approach[EB/OL]. (2019-07-26)[2026-03-30]. https://arxiv.org/abs/1907.11692.[J].arXiv preprint arXiv:1907.11692, 2019, 无(无):无-无
[20]REN S, HE K, GIRSHICK R, et al.Faster R-CNN: Towards real-time object detection with region proposal networks[J].IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 39(6):1137-1149
[21]YANG Z, GONG B, WANG L, et al.A fast and accurate one-stage approach to visual grounding[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Piscataway: IEEE, 2019: 4683-4693..Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, 无(无):4683-4693
[22]YANG Z, CHEN T, WANG L, et al.Improving one-stage visual grounding by recursive sub-query construction[C]//Proceedings of the European Conference on Computer Vision. Cham: Springer, 2020: 387-404..Proceedings of the European Conference on Computer Vision (ECCV), 2020, 无(无):387-404
[23]HUANG B, LIAN D, LUO W, et al.Look before you leap: learning landmark features for one-stage visual grounding[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE, 2021: 16888-16897..Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, 无(无):16888-16897