Learning End-to-End Scene Flow by Distilling Single Tasks Knowledge
Filippo Aleotti, Matteo Poggi, Fabio Tosi, Stefano Mattoccia
Abstract
Scene flow is a challenging task aimed at jointly estimating the 3D structure and motion of the sensed environment. Although deep learning solutions achieve outstanding performance in terms of accuracy, these approaches divide the whole problem into standalone tasks (stereo and optical flow) addressing them with independent networks. Such a strategy dramatically increases the complexity of the training procedure and requires power-hungry GPUs to infer scene flow barely at 1 FPS. Conversely, we propose DWARF, a novel and lightweight architecture able to infer full scene flow jointly reasoning about depth and optical flow easily and elegantly trainable end-to-end from scratch. Moreover, since ground truth images for full scene flow are scarce, we propose to leverage on the knowledge learned by networks specialized in stereo or flow, for which much more data are available, to distill proxy annotations. Exhaustive experiments show that i) DWARF runs at about 10 FPS on a single high-end GPU and about 1 FPS on NVIDIA Jetson TX2 embedded at KITTI resolution, with moderate drop in accuracy compared to 10× deeper models, ii) learning from many distilled samples is more effective than from the few, annotated ones available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b2401ed-0fbc-48d8-9223-7982d6ba778aCited by top-tier papers7
- IFRNet: Intermediate Feature Refine Network for Efficient Frame InterpolationLingtong Kong, Boyuan Jiang, Donghao Luo, Wenqing Chu et al.CVPR 2022 · 166 citations
- CamLiFlow: Bidirectional Camera-LiDAR Fusion for Joint Optical Flow and Scene Flow EstimationHaisong Liu, Tao Lu, Yihui Xu, Jia Liu et al.CVPR 2022 · 64 citations
- Effective Video Abnormal Event Detection by Learning A Consistency-Aware High-Level Feature ExtractorGuang Yu, Siqi Wang, Zhiping Cai, Xinwang Liu et al.ACM MM 2022 · 7 citations
- Self-Supervised Multi-Frame Monocular Scene FlowJunhwa Hur, Stefan RothCVPR 2021
- SMD-Nets: Stereo Mixture Density NetworksFabio Tosi, Yiyi Liao, Carolin Schmitt, Andreas GeigerCVPR 2021
Builds on1
Related papers
- Self-Supervised Monocular Scene Flow EstimationJunhwa Hur, Stefan RothCVPR 2020
- ZeroFlow: Scalable Scene Flow via DistillationKyle Vedder, Neehar Peri, Nathaniel Chodosh, Ishan Khatri et al.ICLR 2024 · 12 citations
- RAFT-3D: Scene Flow Using Rigid-Motion EmbeddingsZachary Teed, Jia DengCVPR 2021
- Flow2Stereo: Effective Self-Supervised Learning of Optical Flow and Stereo MatchingPengpeng Liu, Irwin King, Michael R. Lyu, Jia XuCVPR 2020
- FlowSeek: Optical Flow Made Easier with Depth Foundation Models and Motion BasesMatteo Poggi, Fabio TosiICCV 2025 · 5 citations
