Lite-Mono: A Lightweight CNN and Transformer Architecture for Self-Supervised Monocular Depth Estimation
Ning Zhang, Francesco Nex, George Vosselman, Norman Kerle
摘要
Self-supervised monocular depth estimation that does not require ground truth for training has attracted attention in recent years. It is of high interest to design lightweight but effective models so that they can be deployed on edge devices. Many existing architectures benefit from using heavier backbones at the expense of model sizes. This paper achieves comparable results with a lightweight architecture. Specifically, the efficient combination of CNNs and Transformers is investigated, and a hybrid architecture called Lite-Mono is presented. A Consecutive Dilated Convolutions (CDC) module and a Local-Global Features Interaction (LGFI) module are proposed. The former is used to extract rich multi-scale local features, and the latter takes advantage of the self-attention mechanism to encode longrange global information into the features. Experiments demonstrate that Lite-Mono outperforms Monodepth2 by a large margin in accuracy, with about 80% fewer trainable parameters. Our codes and models are available at https://github.com/noahzn/Lite-Mono.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- Self-supervised Monocular Depth Estimation: Let's Talk About The WeatherKieran Saunders, George Vogiatzis, Luis J. MansoICCV 2023 · 被引用 64 次
- Binocular-Guided 3D Gaussian Splatting with View Consistency for Sparse View SynthesisLiang Han, Junsheng Zhou, Yu-Shen Liu, Zhizhong HanNeurIPS 2024 · 被引用 63 次
- Dynamo-Depth: Fixing Unsupervised Depth Estimation for Dynamical ScenesYihong Sun, Bharath HariharanNeurIPS 2023 · 被引用 58 次
- Jasmine: Harnessing Diffusion Prior for Self-supervised Depth EstimationJiyuan Wang, Chunyu Lin, Cheng Guan, Lang Nie 等NeurIPS 2025 · 被引用 26 次
- Depth-Centric Dehazing and Depth-Estimation from Real-World Hazy Driving VideoJunkai Fan, Kun Wang, Zhiqiang Yan, Xiang Chen 等AAAI 2025 · 被引用 15 次
它引用的顶会 Paper13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 被引用 2,416 次
- CMT: Convolutional Neural Networks Meet Vision TransformersJianyuan Guo, Kai Han, Han Wu, Yehui Tang 等CVPR 2022 · 被引用 839 次
- XCiT: Cross-Covariance Image TransformersAlaaeldin Ali, Hugo Touvron, Mathilde Caron, Piotr Bojanowski 等NeurIPS 2021 · 被引用 692 次
相关 Paper
- Deep Digging into the Generalization of Self-Supervised Monocular Depth EstimationJinwoo Bae, Sungho Moon, Sunghoon ImAAAI 2023 · 被引用 127 次
- Self-Supervised Monocular Trained Depth Estimation Using Self-Attention and Discrete Disparity VolumeAdrian Johnston, Gustavo CarneiroCVPR 2020
- Hybrid-Grained Feature Aggregation with Coarse-to-Fine Language Guidance for Self-Supervised Monocular Depth EstimationWenyao Zhang, Hongsi Liu, Bohan Li, Jiawei He 等ICCV 2025 · 被引用 2 次
- HR-Depth: High Resolution Self-Supervised Monocular Depth EstimationXiaoyang Lyu, Liang Liu, Mengmeng Wang, Xin Kong 等AAAI 2021 · 被引用 341 次
- Transformer-Based Attention Networks for Continuous Pixel-Wise PredictionGuanglei Yang, Hao Tang, Mingli Ding, Nicu Sebe 等ICCV 2021 · 被引用 246 次
