SVDC: Consistent Direct Time-of-Flight Video Depth Completion with Frequency Selective Fusion
Xuan Zhu, Jijun Xiang, Xianqi Wang, Longliang Liu, Yu Wang, Hong Zhang, Fei Guo, Xin Yang
Abstract
Lightweight direct Time-of-Flight (dToF) sensors are ideal for 3D sensing on mobile devices. However, due to the manufacturing constraints of compact devices and the inherent physical principles of imaging, dToF depth maps are sparse and noisy. In this paper, we propose a novel video depth completion method, called SVDC, by fusing the sparse dToF data with the corresponding RGB guidance. Our method employs a multi-frame fusion scheme to mitigate the spatial ambiguity resulting from the sparse dToF imaging. Misalignment between consecutive frames during multi-frame fusion could cause blending between object edges and the background, which results in a loss of detail. To address this, we introduce an adaptive frequency selective fusion (AFSF) module, which automatically selects convolution kernel sizes to fuse multi-frame features. Our AFSF utilizes a channel-spatial enhancement attention (CSEA) module to enhance features and generates an attention map as fusion weights. The AFSF ensures edge detail recovery while suppressing high-frequency noise in smooth regions. To further enhance temporal consistency, We propose a cross-window consistency loss to ensure consistent predictions across different windows, effectively reducing flickering. Our proposed SVDC achieves optimal accuracy and consistency on the TartanAir and Dynamic Replica datasets. Code is available at https://github.com/Lan1eve/SVDC .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- BANet: Bilateral Aggregation Network for Mobile Stereo MatchingGangwei Xu, Jiaxin Liu, Xianqi Wang, Junda Cheng et al.ICCV 2025 · 7 citations
- SpatioTemporal Difference Network for Video Depth Super-ResolutionZhengxue Wang, Yuan Wu, Xiang Li, Zhiqiang Yan et al.AAAI 2026 · 2 citations
- MonSter: Marry Monodepth to Stereo Unleashes PowerJunda Cheng, Longliang Liu, Gangwei Xu, Xianqi Wang et al.CVPR 2025
Builds on24
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and AlignmentKelvin C. K. Chan, Shangchen Zhou, Xiangyu Xu, Chen Change LoyCVPR 2022 · 522 citations
- GMFlow: Learning Optical Flow via Global MatchingHaofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi et al.CVPR 2022 · 353 citations
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen et al.SIGGRAPH 2020 · 321 citations
- CSPN++: Learning Context and Resource Aware Convolutional Spatial Propagation Networks for Depth CompletionXinjing Cheng, Peng Wang, Chenye Guan, Ruigang YangAAAI 2020 · 270 citations
Related papers
- Dense Metric Depth Completion from Sparse Direct Time-of-Flight SensorsHakyeong Kim, Ruicheng Wang, Chengtang Yao, Jiaolong Yang et al.CVPR 2026 · 1 citation
- Consistent Time-of-Flight Depth Denoising via Graph-Informed Geometric AttentionWeida Wang, Changyong He, Jin Zeng, Di QiuICCV 2025 · 1 citation
- Tri-Perspective view Decomposition for Geometry-Aware Depth CompletionZhiqiang Yan, Yuankai Lin, Kun Wang, Yupeng Zheng et al.CVPR 2024 · 33 citations
- Aggregating Feature Point Cloud for Depth CompletionZhu Yu, Zehua Sheng, Zili Zhou, Lun Luo et al.ICCV 2023 · 42 citations
- Completion as Enhancement: A Degradation-Aware Selective Image Guided Network for Depth CompletionZhiqiang Yan, Zhengxue Wang, Kun Wang, Jun Li et al.CVPR 2025
