SceneFun3D: Fine-Grained Functionality and Affordance Understanding in 3D Scenes
Alexandros Delitzas, Ayça Takmaz, Federico Tombari, Robert W. Sumner, Marc Pollefeys, Francis Engelmann
Abstract
Existing 3D scene understanding methods are heavily focused on 3D semantic and instance segmentation. However, identifying objects and their parts only constitutes an intermediate step towards a more fine-grained goal, which is effectively interacting with the functional interactive elements (e.g., handles, knobs, buttons) in the scene to accomplish diverse tasks. To this end, we introduce SceneFun3D, a large-scale dataset with more than 14.8k highly accurate interaction annotations for 710 high-resolution realworld 3D indoor scenes. We accompany the annotations with motion parameter information, describing how to interact with these elements, and a diverse set of natural language descriptions of tasks that involve manipulating them in the scene context. To showcase the value of our dataset, we introduce three novel tasks, namely functionality segmentation, task-driven affordance grounding and 3D motion estimation, and adapt existing state-of-the-art methods to tackle them. Our experiments show that solving these tasks in real 3D scenes remains challenging despite recent progress in closed-set and open-set 3D scene understanding methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers33
- OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel ViewsFrancis Engelmann, Fabian Manhardt, Michael Niemeyer, Keisuke Tateno et al.ICLR 2024 · 69 citations
- EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric ViewsYuhang Yang, Wei Zhai, Chengfeng Wang, Chengjun Yu et al.NeurIPS 2024 · 31 citations
- Learning 2D Invariant Affordance Knowledge for 3D Affordance GroundingXianqiang Gao, Pingrui Zhang, Delin Qu, Dong Wang et al.AAAI 2025 · 20 citations
- AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language ModelsXinyi Wang, Xun Yang, Yanlong Xu, Yuchen Wu et al.NeurIPS 2025 · 18 citations
- Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language ModelsHanqing Wang, Shaoyang Wang, Yiming Zhong, Zemin Yang et al.AAAI 2026 · 13 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui et al.ICCV 2019 · 3,193 citations
- ScanNet++: A High-Fidelity Dataset of 3D Indoor ScenesChandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, Angela DaiICCV 2023 · 659 citations
- LERF: Language Embedded Radiance FieldsJustin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa et al.ICCV 2023 · 620 citations
- OpenMask3D: Open-Vocabulary 3D Instance SegmentationAyça Takmaz, Elisabetta Fedele, Robert W. Sumner, Marc Pollefeys et al.NeurIPS 2023 · 389 citations
Related papers
- Open-Vocabulary Functional 3D Scene Graphs for Real-World Indoor SpacesChenyangguang Zhang, Alexandros Delitzas, Fangjinhua Wang, Ruida Zhang et al.CVPR 2025
- Task-Aware 3D Affordance Segmentation via 2D Guidance and Geometric RefinementLian He, Meng Liu, Qilang Ye, Yu Zhou et al.AAAI 2026 · 3 citations
- AffordMatcher: Affordance Learning in 3D Scenes from Visual SignifiersNghia Vu, Tuong Do, Khang Nguyen, Baoru Huang et al.CVPR 2026 · 2 citations
- Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene DescriptionAnna-Maria Halacheva, Yang Miao, Jan-Nico Zaech, Xi Wang et al.ICCV 2025 · 2 citations
- AffordPose: A Large-scale Dataset of Hand-Object Interactions with Affordance-driven Hand PoseJuntao Jian, Xiuping Liu, Manyi Li, Ruizhen Hu et al.ICCV 2023 · 78 citations
