S2R-DepthNet: Learning a Generalizable Depth-Specific Structural Representation
Xiaotian Chen, Yuwang Wang, Xuejin Chen, Wenjun Zeng
Abstract
Human can infer the 3D geometry of a scene from a sketch instead of a realistic image, which indicates that the spatial structure plays a fundamental role in understanding the depth of scenes. We are the first to explore the learning of a depth-specific structural representation, which captures the essential feature for depth estimation and ignores irrelevant style information. Our S2R-DepthNet (Synthetic to Real DepthNet) can be well generalized to unseen real-world data directly even though it is only trained on synthetic data. S2R-DepthNet consists of: a) a Structure Extraction (STE) module which extracts a domaininvariant structural representation from an image by disentangling the image into domain-invariant structure and domain-specific style components, b) a Depth-specific Attention (DSA) module, which learns task-specific knowledge to suppress depth-irrelevant structures for better depth estimation and generalization, and c) a depth prediction module (DP) to predict depth from the depth-specific representation. Without access of any real-world images, our method even outperforms the state-of-the-art unsupervised domain adaptation methods which use real-world images of the target domain for training. In addition, when using a small amount of labeled real-world data, we achieve the state-ofthe-art performance under the semi-supervised setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e638e64b-ad81-4bc7-b28f-f94f2c6bc16cCited by top-tier papers3
- MORE: A Multimodal Object-Entity Relation Extraction Dataset with a Benchmark EvaluationLiang He, Hongke Wang, Yongchang Cao, Zhen Wu et al.ACM MM 2023 · 17 citations
- Dual-branch Graph Feature Learning for NLOS ImagingXiongfei Su, Tianyi Zhu, Lina Liu, Zheng Chen et al.AAAI 2025 · 4 citations
- FreeDNA: Endowing Domain Adaptation of Diffusion-Based Dense Prediction with Training-Free Domain Noise AlignmentHang Xu, Jie Huang, Linjiang Huang, Dong Liu et al.ICCV 2025 · 1 citation
Builds on11
- Enforcing Geometric Constraints of Virtual Normal for Depth PredictionWei Yin, Yifan Liu, Chunhua Shen, Youliang YanICCV 2019 · 487 citations
- Self-Supervised Monocular Depth HintsJamie Watson, Michael Firman, Gabriel J. Brostow, Daniyar TurmukhambetovICCV 2019 · 287 citations
- Self-Supervised Learning With Geometric Constraints in Monocular Video: Connecting Flow, Depth, and CameraYuhua Chen, Cordelia Schmid, Cristian SminchisescuICCV 2019 · 265 citations
- Visualization of Convolutional Neural Networks for Monocular Depth EstimationJunjie Hu, Yan Zhang, Takayuki OkataniICCV 2019 · 91 citations
- Unsupervised High-Resolution Depth Learning From Videos With Dual NetworksJunsheng Zhou, Yuwang Wang, Kaihuai Qin, Wenjun ZengICCV 2019 · 77 citations
Related papers
- Semantically-Guided Representation Learning for Self-Supervised Monocular DepthVitor Guizilini, Rui Hou, Jie Li, Rares Ambrus et al.ICLR 2020 · 264 citations
- Geometry-Aware Network for Domain Adaptive Semantic SegmentationYinghong Liao, Wending Zhou, Xu Yan, Zhen Li et al.AAAI 2023 · 9 citations
- DADA: Depth-Aware Domain Adaptation in Semantic SegmentationTuan-Hung Vu, Himalaya Jain, Maxime Bucher, Matthieu Cord et al.ICCV 2019 · 202 citations
- Transferring to Real-World Layouts: A Depth-aware Framework for Scene AdaptationMu Chen, Zhedong Zheng, Yi YangACM MM 2024 · 19 citations
- DRANet: Disentangling Representation and Adaptation Networks for Unsupervised Cross-Domain AdaptationSeunghun Lee, Sunghyun Cho, Sunghoon ImCVPR 2021
