Rel3D: A Minimally Contrastive Benchmark for Grounding Spatial Relations in 3D
Ankit Goyal, Kaiyu Yang, Dawei Yang, Jia Deng
摘要
Understanding spatial relations (e.g., "laptop on table") in visual input is important for both humans and robots. Existing datasets are insufficient as they lack large-scale, high-quality 3D ground truth information, which is critical for learning spatial relations. In this paper, we fill this gap by constructing Rel3D: the first large-scale, human-annotated dataset for grounding spatial relations in 3D. Rel3D enables quantifying the effectiveness of 3D information in predicting spatial relations on large-scale human data. Moreover, we propose minimally contrastive data collection -- a novel crowdsourcing method for reducing dataset bias. The 3D scenes in our dataset come in minimally contrastive pairs: two scenes in a pair are almost identical, but a spatial relation holds in one and fails in the other. We empirically validate that minimally contrastive examples can diagnose issues with current relation detection models as well as lead to sample-efficient training. Code and data are available at https://github.com/princeton-vl/Rel3D.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- IFOR: Iterative Flow Minimization for Robotic Object RearrangementAnkit Goyal, Arsalan Mousavian, Chris Paxton, Yu-Wei Chao 等CVPR 2022 · 被引用 34 次
- Can Transformers Capture Spatial Relations between Objects?Chuan Wen, Dinesh Jayaraman, Yang GaoICLR 2024 · 被引用 11 次
- MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMsErik A. Daxberger, Nina Wenzel, David Griffiths, Haiming Gang 等ICCV 2025 · 被引用 10 次
- Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language ModelsYuan-Hong Liao, Rafid Mahmood, Sanja Fidler, David AcunaEMNLP 2024 · 被引用 4 次
- ShapeTalk: A Language Dataset and Framework for 3D Shape Edits and DeformationsPanos Achlioptas, Ian Huang, Minhyuk Sung, Sergey Tulyakov 等CVPR 2023
它引用的顶会 Paper2
相关 Paper
- Multi3DRefer: Grounding Text Description to Multiple 3D ObjectsYiming Zhang, ZeMing Gong, Angel X. ChangICCV 2023 · 被引用 157 次
- AugRefer: Advancing 3D Visual Grounding via Cross-Modal Augmentation and Spatial Relation-based ReferringXinyi Wang, Na Zhao, Zhiyuan Han, Dan Guo 等AAAI 2025 · 被引用 12 次
- Learning Multi-View Spatial Reasoning from Cross-View RelationsSuchae Jeong, Jaehwi Song, Haeone Lee, Hanna Kim 等CVPR 2026
- CHORUS: Learning Canonicalized 3D Human-Object Spatial Relations from Unbounded Synthesized ImagesSookwan Han, Hanbyul JooICCV 2023 · 被引用 19 次
- AmbiRefer3D: 3D Visual Grounding with Referential AmbiguityRongjiang Zhu, Wei Kang, Zeqi Liu, Chen junyu 等ICML 2026
