Lightweight monocular depth estimation model based on multi-kernel grouped convolution and adaptive feature fusion
Liu Yao1, Huang He1, Ma Yanjie1, Ma Chaowei2, Ren Boyang1, Yang Junxing1
1. School of Geomatics and Urban Spatial Information, Beijing University of Civil Engineering and Architecture, Beijing 102616, China; 2. China Aerospace Planning and Design Group Co., Ltd., Beijing 100070, China
Liu Yao, Huang He, Ma Yanjie, Ma Chaowei, Ren Boyang, Yang Junxing. Lightweight monocular depth estimation model based on multi-kernel grouped convolution and adaptive feature fusion[J]. Bulletin of Surveying and Mapping, 2026, 0(8): 67-73,81.
[1] 胡海洋,陈超平,高天沐,等.单/双目深度估计研究进展与应用综述[J].红外与激光工程,2025,54(7):25-38. [2] 郭迟,刘阳,罗亚荣,等.图像语义信息在视觉SLAM中的应用研究进展[J].测绘学报,2024,53(6):1057-1076. [3] Ranjan A,Jampani V,Balles L,et al.Competitive collaboration:joint unsupervised learning of depth,camera motion,optical flow and motion segmentation[C]//Proceedings of 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).Long Beach:IEEE,2020:12232-12241. [4] Wang Chaoyang,Buenaposada J M,Zhu Rui,et al.Learning depth from monocular videos using direct methods[C]//Proceedings of 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition.Salt Lake City:IEEE,2018:2022-2030. [5] 蒋祥龙,邓文亮,何胜喜.单目视觉驱动的机器人实时高精度稠密场景重建算法[J].测绘通报,2025(10):71-75. [6] Casser V,Pirk S,Mahjourian R,et al.Depth prediction without the sensors:leveraging structure for unsupervised learning from monocular videos[J].Proceedings of the AAAI Conference on Artificial Intelligence,2019,33(1):8001-8008. [7] Sun Libo,Bian Jiawang,Zhan Huangying,et al.SC-DepthV3:robust self-supervised monocular depth estimation for dynamic scenes[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2024,46(1):497-508. [8] Zhou Zhongkai,Fan Xinnan,Shi Pengfei,et al.R-MSFM:recurrent multi-scale feature modulation for monocular depth estimating[C]//Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision.Montreal:IEEE,2022:12757-12766. [9] Zou Yuliang,Luo Zelun,Huang Jiabin.DF-Net:unsupervised joint learning of depth and flow using cross-task consistency[C]//Proceedings of 2018 Computer Vision-ECCV.Cham:Springer,2018:38-55. [10] Luo Chenxu,Yang Zhenheng,Wang Peng,et al.Every pixel counts:joint learning of geometry and motion with 3D holistic understanding[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2020,42(10):2624-2641. [11] Godard C,Mac Aodha O,Firman M,et al.Digging into self-supervised monocular depth estimation[C]//Proceedings of 2019 IEEE/CVF International Conference on Computer Vision.Seoul:IEEE,2019:3827-3837. [12] Han D,Shin J,Kim N,et al.TransDSSL:transformer based depth estimation via self-supervised learning[J].IEEE Robotics and Automation Letters,2022,7(4):10969-10976. [13] Zhang Ning,Nex F,Vosselman G,et al.Lite-mono:a lightweight CNN and transformer architecture for self-supervised monocular depth estimation[C]//Proceedings of 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition.Vancouver:IEEE,2023:18537-18546. [14] Zhang Jiacheng,Wang Xichao,Wang Xinyao,et al.CT-mono:leveraging CNNs and transformers for self-supervised depth estimation in single-view scenarios[J].Arabian Journal for Science and Engineering,2026,51(5):6023-6038. [15] Bae J,Moon S,Im S.Deep digging into the generalization of self-supervised monocular depth estimation[J].Proceedings of the AAAI Conference on Artificial Intelligence,2023,37(1):187-196. [16] Dai Yimian,Gieseke F,Oehmcke S,et al.Attentional feature fusion[C]//Proceedings of 2021 IEEE Winter Conference on Applications of Computer Vision.Waikoloa:IEEE,2021:3559-3568. [17] Eigen D,Fergus R.Predicting depth,surface normals and semantic labels with a common multi-scale convolutional architecture[C]//Proceedings of 2015 IEEE International Conference on Computer Vision.Santiago:IEEE,2016:2650-2658.