Bulletin of Surveying and Mapping ›› 2026, Vol. 0 ›› Issue (8): 67-73,81.doi: 10.13474/j.cnki.11-2246.2026.0810

Previous Articles    

Lightweight monocular depth estimation model based on multi-kernel grouped convolution and adaptive feature fusion

Liu Yao1, Huang He1, Ma Yanjie1, Ma Chaowei2, Ren Boyang1, Yang Junxing1   

  1. 1. School of Geomatics and Urban Spatial Information, Beijing University of Civil Engineering and Architecture, Beijing 102616, China;
    2. China Aerospace Planning and Design Group Co., Ltd., Beijing 100070, China
  • Received:2025-12-08 Published:2026-09-12

Abstract: [Purposes] To address the problems of insufficient extraction of distant object edge details,weak global consistency of depth prediction,and the difficulty of meeting real-time 3D perception requirements in autonomous driving for self-supervised monocular depth estimation,this paper proposes a lightweight self-supervised monocular depth estimation model. [Methods] Firstly,a multi-kernel grouped depthwise separable convolution module is designed to effectively fuse local detail and global structural information,significantly enhancing feature representation capability.In addition,an attention-guided adaptive feature fusion module is introduced to dynamically adjust multi-scale feature weighting and enhance the overall consistency of depth prediction. [Findings] Experimental results show that,compared with classical methods Monodepth2,the proposed approach reduces parameters by 10.4% and computation by 34.7%,achieving 0.112 and 0.878 on key performance metrics Abs Rel and δ1,respectively,and demonstrating superior generalization performance in cross-scene testing. [Conclusions] The proposed method effectively improves depth estimation accuracy under a lightweight design,providing a reliable solution for real-time 3D perception in autonomous driving.

Key words: monocular depth estimation, autonomous driving, self-supervised learning, adaptive feature fusion, multi-kernel grouped convolution, depth separable

CLC Number: