导航

Acta Aeronautica et Astronautica Sinica

Previous Articles     Next Articles

Spatial Relation-Driven Remote Sensing Visual Grounding Method

  

  • Received:2026-04-02 Revised:2026-07-20 Online:2026-07-22 Published:2026-07-22

Abstract: Remote sensing visual grounding aims to precisely locate target ground objects in images via natural language queries, serving as a critical technology for intelligent interpretation of remote sensing images. Existing remote sensing visual grounding methods suffer from two core bottlenecks. First, coarse spatial relation annotations and monotonous semantics in training data hinder models from learning robust spatial reasoning capabilities from weakly annotated texts. Second, explicit modeling of primary and secondary semantic relationships within complex queries is absent; in scenarios with redundant semantics or severe reference object interference, the model tends to misclassify distractors as localization targets. To address the above issues, this paper proposes a spatial relation-driven remote sensing visual grounding method named SRMVG. First, a DIOR-RSVG-SR dataset with elaborate spatial relation annotations is constructed to provide high-quality spatial semantic supervision signals. Second, at the cross-modal reasoning stage, a spatial relation alignment framework based on two-branch graph neural networks is proposed. It realizes text-image alignment through structured parsing of four core elements from text features, and introduces a geographic orientation constraint mask to rectify mismatched spatial semantics. Finally, a semantic priority-guided filtering module is designed for fine-grained target localization, which suppresses target misjudgments caused by reference object interference via an adaptive weight fusion mechanism. Experimental results demonstrate that the proposed method achieves overall performance superior to state-of-the-art approaches on two public datasets, DIOR-RSVG and OPT-RSVG. On the self-built DIOR-RSVG-SR dataset, SRMVG achieves remarkably greater performance gains compared with other competitive methods, which verifies the superiority of the proposed method in leveraging elaborate spatial supervision signals.

Key words: remote sensing image interpretation, visual grounding, spatial relation modeling, cross-modal alignment, multi-scale feature fusion

CLC Number: