QuartDepth: Post-Training Quantization for Real-Time Depth Estimation on the Edge
Xuan Shen, Weize Ma, Jing Liu, Changdi Yang, Rui Ding, Quanyi Wang, Henghui Ding, Wei Niu, Yanzhi Wang, Pu Zhao, Jun Lin, Jiuxiang Gu
Abstract
Monocular Depth Estimation (MDE) has emerged as a pivotal task in computer vision, supporting numerous realworld applications. However, deploying accurate depth estimation models on resource-limited edge devices, especially Application-Specific Integrated Circuits (ASICs), is challenging due to the high computational and memory demands. Recent advancements in foundational depth estimation deliver impressive results but further amplify the difficulty of deployment on ASICs. To address this, we propose Quart-Depth which adopts post-training quantization to quantize MDE models with hardware accelerations for ASICs. Our approach involves quantizing both weights and activations to 4-bit precision, reducing the model size and computation cost. To mitigate the performance degradation, we introduce activation polishing and compensation algorithm applied before and after activation quantization, as well as a weight reconstruction method for minimizing errors in weight quantization. Furthermore, we design a flexible and programmable hardware accelerator by supporting kernel fusion and customized instruction programmability, enhancing throughput and efficiency. Experimental results demonstrate that our framework achieves competitive accuracy while enabling fast inference and higher energy efficiency on ASICs, bridging the gap between high-performance depth estimation and practical edge-device applicability. Code: https: //github.com/shawnricecake/quart-depth
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7e068ae-6201-4e18-b5b2-6a51ac34c500Cited by top-tier papers2
- Efficient Reasoning with Hidden ThinkingXuan Shen, Yizhou Wang, Yufa Zhou, Xiangxi Shi et al.ICML 2026 · 56 citations
- LoPrune: Efficient Data Pruning for LoRA-Based Fine-Tuning of Vision TransformerQiang He, Yaozong Yang, KAIBIN WANG, Ziteng Wei et al.CVPR 2026
Builds on43
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos et al.ICML 2020 · 816 citations
Related papers
- QSCA: Quantization with Self-Compensating Auxiliary for Monocular Depth EstimationJincheol Yang, Jaemin Choi, Matti Zinke, Suk-Ju KangNeurIPS 2025 · 1 citation
- Point4Bit: Post Training 4-bit Quantization for Point Cloud 3D DetectionJianyu Wang, Yu Wang, Shengjie Zhao, Sifan ZhouNeurIPS 2025 · 2 citations
- Predicting High-precision Depth on Low-Precision Devices Using 2D Hilbert CurvesMykhailo Uss, Ruslan Yermolenko, Oleksii Shashko, Olena Kolodiazhna et al.ICML 2025
- A2Q: Accumulator-Aware Quantization with Guaranteed Overflow AvoidanceIan Colbert, Alessandro Pappalardo, Jakoba Petri-KoenigICCV 2023 · 19 citations
- Octo: INT8 Training with Loss-aware Compensation and Backward Quantization for Tiny On-device LearningQihua Zhou, Song Guo, Zhihao Qu, Jingcai Guo et al.USENIX ATC 2021 · 55 citations
