Instance-wise Occlusion and Depth Orders in Natural Scenes
Hyunmin Lee, Jaesik Park
Abstract
In this paper, we introduce a new dataset, named In-staOrder, that can be used to understand the geometrical relationships of instances in an image. The dataset consists of 2.9M annotations of geometric orderings for class-labeled instances in 101K natural scenes. The scenes were annotated by 3,659 crowd-workers regarding (1) occlusion order that identifies occluder/occludee and (2) depth order that describes ordinal relations that consider relative distance from the camera. The dataset provides joint annotation of two kinds of orderings for the same instances, and we discover that the occlusion order and depth order are complementary. We also introduce a geometric order prediction network called InstaOrderNet, which is superior to state-of-the-art approaches. Moreover, we propose a dense depth prediction network called InstaDepthNet that uses auxiliary geometric order loss to boost the accuracy of the state-of-the-art depth prediction approach, MiDaS [54].
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 656186ad-ad93-4baa-8d40-e626d261d702Cited by top-tier papers14
- Amodal Completion via Progressive Mixed Context DiffusionKatherine Xu, Lingzhi Zhang, Jianbo ShiCVPR 2024 · 20 citations
- Occ2Net: Robust Image Matching Based on 3D Occupancy Estimation for Occluded RegionsMiao Fan, Mingrui Chen, Chen Hu, Shuchang ZhouICCV 2023 · 7 citations
- MULAN: A Multi Layer Annotated Dataset for Controllable Text-to-Image GenerationPetru-Daniel Tudosiu, Yongxin Yang, Shifeng Zhang, Fei Chen et al.CVPR 2024 · 7 citations
- I2E: From Image Pixels to Actionable Interactive Environments for Text-Guided Image EditingJinghan Yu, Junhao Xiao, Chenyu Zhu, Jiaming Li et al.ACL 2026 · 3 citations
- SynergyAmodal: Deocclude Anything with Text ControlXinyang Li, Chengjie Yi, Jiawei Lai, Mingbao Lin et al.ACM MM 2025 · 3 citations
Builds on18
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- DeepV2D: Video to Depth with Differentiable Structure from MotionZachary Teed, Jia DengICLR 2020 · 314 citations
- Specifying Object Attributes and Relations in Interactive Scene GenerationOron Ashual, Lior WolfICCV 2019 · 190 citations
- Exploiting Temporal Consistency for Real-Time Video Depth EstimationHaokui Zhang, Ying Li, Yuanzhouhan Cao, Yu Liu et al.ICCV 2019 · 137 citations
- Visualizing the Invisible: Occluded Vehicle Segmentation and RecoveryXiaosheng Yan, Yuanlong Yu, Feigege Wang, Wenxi Liu et al.ICCV 2019 · 46 citations
Related papers
- Holistic Order Prediction in Natural ScenesPierre Musacchio, Hyunmin Lee, Jaesik ParkNeurIPS 2025 · 1 citation
- Instance-Level Video Depth in Groups Beyond OcclusionsYuan Liang, Yang Zhou, Ziming Sun, Tianyi Xiang et al.ICCV 2025
- Order-aware Human Interaction ManipulationMandi Luo, Jie Cao, Ran HeACM MM 2022 · 1 citation
- Robust Instance Segmentation Through Reasoning About Multi-Object OcclusionXiaoding Yuan, Adam Kortylewski, Yihong Sun, Alan L. YuilleCVPR 2021
- Revealing the Reciprocal Relations between Self-Supervised Stereo and Monocular Depth EstimationZhi Chen, Xiaoqing Ye, Wei Yang, Zhenbo Xu et al.ICCV 2021 · 34 citations
