测绘通报 ›› 2026, Vol. 0 ›› Issue (8): 67-73,81.doi: 10.13474/j.cnki.11-2246.2026.0810

• 学术研究 • 上一篇    

融合多核分组卷积与自适应特征的轻量化单目深度估计模型

刘瑶1, 黄鹤1, 马延杰1, 马朝威2, 任博洋1, 杨军星1   

  1. 1. 北京建筑大学测绘与城市信息空间学院, 北京 102616;
    2. 航天规划设计集团有限公司, 北京 100070
  • 收稿日期:2025-12-08 发布日期:2026-09-12
  • 通讯作者: 杨军星。E-mail:yangjunxing@bucea.edu.cn
  • 作者简介:刘瑶(2001—),女,硕士生,主要从事深度估计、计算机视觉等相关研究。E-mail:2108570424065@stu.bucea.edu.cn
  • 基金资助:
    国家自然科学基金(42201483); 北京建筑大学青年教师科研能力提升计划(X23001)

Lightweight monocular depth estimation model based on multi-kernel grouped convolution and adaptive feature fusion

Liu Yao1, Huang He1, Ma Yanjie1, Ma Chaowei2, Ren Boyang1, Yang Junxing1   

  1. 1. School of Geomatics and Urban Spatial Information, Beijing University of Civil Engineering and Architecture, Beijing 102616, China;
    2. China Aerospace Planning and Design Group Co., Ltd., Beijing 100070, China
  • Received:2025-12-08 Published:2026-09-12

摘要: [目的] 针对自监督单目深度估计远距离物体边缘细节提取不足、深度整体一致性表达较弱及难以满足自动驾驶对实时三维感知需求等问题,本文提出了一种轻量化自监督单目深度估计模型。[方法] 首先设计了多核分组深度可分离卷积模块,实现局部细节与全局结构信息的有效融合,大幅增强特征表达能力;然后提出了一种注意力引导的自适应特征融合模块,以动态调整多尺度特征的权重分配,增强深度预测的整体一致性。[结果] 试验结果表明,与经典算法Monodepth2相比,本文方法参数量减少10.4%、计算量减少34.7%,在关键性能指标Abs Rel和δ1上分别达到0.112和0.878,并在跨场景测试中展现出更优的泛化性能。[结论] 本文方法在轻量化设计下能有效提升深度估计精度,为自动驾驶实时三维感知提供了可靠方案。

关键词: 单目深度估计, 自动驾驶, 自监督, 自适应特征融合, 多核分组卷积, 深度可分离

Abstract: [Purposes] To address the problems of insufficient extraction of distant object edge details,weak global consistency of depth prediction,and the difficulty of meeting real-time 3D perception requirements in autonomous driving for self-supervised monocular depth estimation,this paper proposes a lightweight self-supervised monocular depth estimation model. [Methods] Firstly,a multi-kernel grouped depthwise separable convolution module is designed to effectively fuse local detail and global structural information,significantly enhancing feature representation capability.In addition,an attention-guided adaptive feature fusion module is introduced to dynamically adjust multi-scale feature weighting and enhance the overall consistency of depth prediction. [Findings] Experimental results show that,compared with classical methods Monodepth2,the proposed approach reduces parameters by 10.4% and computation by 34.7%,achieving 0.112 and 0.878 on key performance metrics Abs Rel and δ1,respectively,and demonstrating superior generalization performance in cross-scene testing. [Conclusions] The proposed method effectively improves depth estimation accuracy under a lightweight design,providing a reliable solution for real-time 3D perception in autonomous driving.

Key words: monocular depth estimation, autonomous driving, self-supervised learning, adaptive feature fusion, multi-kernel grouped convolution, depth separable

中图分类号: