测绘通报 ›› 2026, Vol. 0 ›› Issue (8): 35-43.doi: 10.13474/j.cnki.11-2246.2026.0806

• 学术研究 • 上一篇    下一篇

掩码先验与区域语义全局融合的遥感影像语义分割方法

荣慧娟1,2, 翟亮1,2, 刘振东2, 陈鑫祥3, 傅毓4, 孙云川2, 李敏3, 何晓辉4   

  1. 1. 兰州交通大学测绘与地理信息学院, 甘肃 兰州 730000;
    2. 中国测绘科学研究院, 北京 100036;
    3. 广东省国土资源技术中心, 广东 广州 510000;
    4. 广西壮族自治区地理信息测绘院, 广西 南宁 530000
  • 收稿日期:2025-12-01 发布日期:2026-09-12
  • 通讯作者: 翟亮。E-mail:zhailiang@casm.ac.cn
  • 作者简介:荣慧娟(2000—),女,硕士生,主要研究方向为实景三维构建。E-mail:15192452621@163.com
  • 基金资助:
    2024年度自然资源部部省合作项目(2024ZRBSHZ172;2024ZRBSHZ155);中央级科研院所基本科研业务费(AR2413)

Semantic segmentation method for remote sensing image through global fusion of mask prior and regional semantics

Rong Huijuan1,2, Zhai Liang1,2, Liu Zhendong2, Chen Xinxiang3, Fu Yu4, Sun Yunchuan2, Li Min3, He Xiaohui4   

  1. 1. College of Geomatics and Geographical Information, Lanzhou Jiaotong University, Lanzhou 730000, China;
    2. Chinese Academy of Surveying and Mapping, Beijing 100036, China;
    3. Guangdong Land and Resources Technology Center, Guangzhou 510000, China;
    4. Guangxi Zhuang Autonomous Region Geographic Information and Surveying Institute, Nanning 530000, China
  • Received:2025-12-01 Published:2026-09-12

摘要: [目的] 针对高分辨率遥感影像语义分割中地物边界模糊、小目标碎片化及区域语义不一致的问题,本文提出一种顾及边界先验与区域语义一致性的遥感影像分割融合方法。[方法] 该方法以分割万物模型SAM生成的无类别几何掩码作为边界先验,与UNetFormer的语义概率在栅格上对齐;在连通域层面构建区域语义特征与代价矩阵,将区域-类别分配形式化为二进制线性整数规划问题并通过全局一致性约束求解。[结果] 在UAVid等3个遥感数据集上的试验表明,本文方法在平均交并比(mIoU)和平均F1值(mF1)上均显著优于单一语义分割模型;尤其在建筑物和车辆等目标上显著改善了边界完整性与语义一致性。[结论] 该方法与主干网络解耦且无需重新训练,可作为独立语义融合模块嵌入现有分割流程,为遥感影像精细语义解译提供了有效途径。

关键词: 高分辨率遥感影像, 语义分割, 分割万物模型, 区域语义聚合, 二进制线性整数规划, 全局一致性约束

Abstract: [Purposes] To address blurred object boundaries,fragmented small targets,and regional semantic inconsistencies commonly encountered in high-resolution remote sensing image segmentation,this study introduces a fusion method that incorporates boundary priors with region-level semantic consistency. [Methods] The proposed approach employs class-agnostic geometric masks generated by SAM(segment anything model)as boundary priors and aligns them with pixel-level semantic probabilities derived from UNetFormer on a unified raster grid.At the connected-component level,a region-semantic feature cost matrix is constructed,and the region-to-class assignment is formulated as a binary linear integer programming problem solved under global semantic consistency constraints. [Findings] Experiments on three remote sensing datasets,including UAVid,show that the proposed method significantly outperforms single-model baselines in mean intersection-over-union (mIoU)and mean F1 (mF1),with notable gains in boundary completeness and semantic consistency for classes such as buildings and vehicles. [Conclusions] The method is decoupled from backbone architectures and requires no retraining,enabling its integration as an independent semantic fusion module within existing segmentation pipelines.It offers an effective solution for enhancing fine-grained semantic interpretation in high-resolution remote sensing imagery.

Key words: high-resolution remote sensing image, semantic segmentation, segment anything model, region semantic aggregation, binary integer linear programming, global consistency constraints

中图分类号: