首页 >

基于空间关系驱动的遥感视觉定位方法-AI+空天科学

刘崇基1,张一鸣2,刘瑜2,姜智卓3,毛永强2,张霖平2,何友4   

  1. 1. 清华大学深圳国际研究生院
    2. 清华大学电子工程系
    3. 南开大学
    4. 海军航空大学
  • 收稿日期:2026-04-02 修回日期:2026-07-20 出版日期:2026-07-22 发布日期:2026-07-22
  • 通讯作者: 姜智卓
  • 基金资助:
    博士后基金;博士后创新人才支持;国家自然科学基金项目;国家自然科学基金项目;国家自然科学基金项目;国家自然科学基金项目

Spatial Relation-Driven Remote Sensing Visual Grounding Method

  • Received:2026-04-02 Revised:2026-07-20 Online:2026-07-22 Published:2026-07-22

摘要: 遥感视觉定位旨在通过自然语言查询实现影像中目标地物的精准定位,是遥感影像智能解译的关键技术。现有遥感视觉定位方法面临两类核心瓶颈:一是训练数据中空间关系标注粗糙且语义单一,导致模型难以从弱标注文本中习得可靠的空间推理能力;二是缺乏对复杂查询中语义主次关系的显式建模,在语义冗杂或参照物干扰显著的场景下,模型易将干扰项误判为定位主体。针对上述问题,本文提出基于空间关系驱动的遥感视觉定位方法SRMVG。首先,构建了包含精细空间关系标注的DIOR-RSVG-SR数据集,提供高质量的空间语义监督信号;其次,在跨模态推理层面提出基于双分支图神经网络的空间关系对齐框架,通过对文本特征进行四要素结构化解析实现图文对齐,并引入地理方位约束掩码以校正空间语义错配;最后,在目标精细定位方面提出语义优先级引导的筛选模块,通过自适应权重融合机制抑制参照物干扰引发的目标误判。实验结果表明,所提方法在DIOR-RSVG和OPT-RSVG两个公开数据集上均取得了优于现有方法的综合性能;在构建的DIOR-RSVG-SR数据集上,SRMVG的提升幅度显著优于其他对比方法,验证了本文方法在利用精细空间监督信号方面的优越性。

关键词: 遥感视觉定位, 空间关系建模, 图神经网络, 跨模态对齐, 遥感图像解译

Abstract: Remote sensing visual grounding aims to precisely locate target ground objects in images via natural language queries, serving as a critical technology for intelligent interpretation of remote sensing images. Existing remote sensing visual grounding methods suffer from two core bottlenecks. First, coarse spatial relation annotations and monotonous semantics in training data hinder models from learning robust spatial reasoning capabilities from weakly annotated texts. Second, explicit modeling of primary and secondary semantic relationships within complex queries is absent; in scenarios with redundant semantics or severe reference object interference, the model tends to misclassify distractors as localization targets. To address the above issues, this paper proposes a spatial relation-driven remote sensing visual grounding method named SRMVG. First, a DIOR-RSVG-SR dataset with elaborate spatial relation annotations is constructed to provide high-quality spatial semantic supervision signals. Second, at the cross-modal reasoning stage, a spatial relation alignment framework based on two-branch graph neural networks is proposed. It realizes text-image alignment through structured parsing of four core elements from text features, and introduces a geographic orientation constraint mask to rectify mismatched spatial semantics. Finally, a semantic priority-guided filtering module is designed for fine-grained target localization, which suppresses target misjudgments caused by reference object interference via an adaptive weight fusion mechanism. Experimental results demonstrate that the proposed method achieves overall performance superior to state-of-the-art approaches on two public datasets, DIOR-RSVG and OPT-RSVG. On the self-built DIOR-RSVG-SR dataset, SRMVG achieves remarkably greater performance gains compared with other competitive methods, which verifies the superiority of the proposed method in leveraging elaborate spatial supervision signals.

Key words: remote sensing image interpretation, visual grounding, spatial relation modeling, cross-modal alignment, multi-scale feature fusion

中图分类号: