Functionality Understanding and Segmentation in 3D Scenes
Jaime Corsetti, Francesco Giuliari, Alice Fasoli, Davide Boscaini, Fabio Poiesi
Abstract
Turn on the ceiling light using the switch next to the TV Open the top right drawer of the cabinet with the TV on top Fun3DU #4f1787ff #eb3678ff #fb773cff Open the second drawer of the cabinet to the left of the TV #bae1ff World Knowledge & Vision Perception Figure 1 . We present Fun3DU, the first method for functionality understanding and segmentation in 3D scenes. Fun3DU interprets natural language descriptions (left-hand side) in order to segment functional objects in real-world 3D environments (right-hand side). Fun3DU relies on world knowledge and vision perception capabilities of pre-trained vision and language models, without requiring task-specific finetuning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9730281f-a26e-451c-b48b-0c040fd3a37dCited by top-tier papers2
- AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language ModelsXinyi Wang, Xun Yang, Yanlong Xu, Yuchen Wu et al.NeurIPS 2025 · 18 citations
- Task-Aware 3D Affordance Segmentation via 2D Guidance and Geometric RefinementLian He, Meng Liu, Qilang Ye, Yu Zhou et al.AAAI 2026 · 3 citations
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- LERF: Language Embedded Radiance FieldsJustin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa et al.ICCV 2023 · 620 citations
Related papers
- SceneFun3D: Fine-Grained Functionality and Affordance Understanding in 3D ScenesAlexandros Delitzas, Ayça Takmaz, Federico Tombari, Robert W. Sumner et al.CVPR 2024
- Open-Vocabulary Functional 3D Scene Graphs for Real-World Indoor SpacesChenyangguang Zhang, Alexandros Delitzas, Fangjinhua Wang, Ruida Zhang et al.CVPR 2025
- FUN REC * Reconstructing Functional 3D Scenes from Egocentric Interaction VideosAlexandros Delitzas, Chenyangguang Zhang, Alexey Gavryushin, Tommaso Di Mario et al.CVPR 2026
- IFR-Explore: Learning Inter-object Functional Relationships in 3D Indoor ScenesQi Li, Kaichun Mo, Yanchao Yang, Hang Zhao et al.ICLR 2022 · 9 citations
- TVDRNet: Text-driven Viewpoint Optimization via Differentiable Rendering for 3D Reasoning SegmentationTingran Wang, Changshuo Wang, Pinjie Xu, ZhangHuang et al.ICML 2026
