Rel3D: A Minimally Contrastive Benchmark for Grounding Spatial Relations in 3D
Ankit Goyal, Kaiyu Yang, Dawei Yang, Jia Deng
Abstract
Understanding spatial relations (e.g., "laptop on table") in visual input is important for both humans and robots. Existing datasets are insufficient as they lack large-scale, high-quality 3D ground truth information, which is critical for learning spatial relations. In this paper, we fill this gap by constructing Rel3D: the first large-scale, human-annotated dataset for grounding spatial relations in 3D. Rel3D enables quantifying the effectiveness of 3D information in predicting spatial relations on large-scale human data. Moreover, we propose minimally contrastive data collection -- a novel crowdsourcing method for reducing dataset bias. The 3D scenes in our dataset come in minimally contrastive pairs: two scenes in a pair are almost identical, but a spatial relation holds in one and fails in the other. We empirically validate that minimally contrastive examples can diagnose issues with current relation detection models as well as lead to sample-efficient training. Code and data are available at https://github.com/princeton-vl/Rel3D.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 447773b1-7b77-4f8b-86d6-bcf6df5ba232Cited by top-tier papers6
- IFOR: Iterative Flow Minimization for Robotic Object RearrangementAnkit Goyal, Arsalan Mousavian, Chris Paxton, Yu-Wei Chao et al.CVPR 2022 · 34 citations
- Can Transformers Capture Spatial Relations between Objects?Chuan Wen, Dinesh Jayaraman, Yang GaoICLR 2024 · 11 citations
- MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMsErik A. Daxberger, Nina Wenzel, David Griffiths, Haiming Gang et al.ICCV 2025 · 10 citations
- Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language ModelsYuan-Hong Liao, Rafid Mahmood, Sanja Fidler, David AcunaEMNLP 2024 · 4 citations
- ShapeTalk: A Language Dataset and Framework for 3D Shape Edits and DeformationsPanos Achlioptas, Ian Huang, Minhyuk Sung, Sergey Tulyakov et al.CVPR 2023
Builds on2
Related papers
- Multi3DRefer: Grounding Text Description to Multiple 3D ObjectsYiming Zhang, ZeMing Gong, Angel X. ChangICCV 2023 · 157 citations
- AugRefer: Advancing 3D Visual Grounding via Cross-Modal Augmentation and Spatial Relation-based ReferringXinyi Wang, Na Zhao, Zhiyuan Han, Dan Guo et al.AAAI 2025 · 12 citations
- Learning Multi-View Spatial Reasoning from Cross-View RelationsSuchae Jeong, Jaehwi Song, Haeone Lee, Hanna Kim et al.CVPR 2026
- CHORUS: Learning Canonicalized 3D Human-Object Spatial Relations from Unbounded Synthesized ImagesSookwan Han, Hanbyul JooICCV 2023 · 19 citations
- AmbiRefer3D: 3D Visual Grounding with Referential AmbiguityRongjiang Zhu, Wei Kang, Zeqi Liu, Chen junyu et al.ICML 2026
