InstructPart: Task-Oriented Part Segmentation with Instruction Reasoning
Zifu Wan, Yaqi Xie, Ce Zhang, Zhiqiu Lin, Zihan Wang, Simon Stepputtis, Deva Ramanan, Katia P. Sycara
Abstract
Large multimodal foundation models, particularly in the domains of language and vision, have significantly advanced various tasks, including robotics, autonomous driving, information retrieval, and grounding. However, many of these models perceive objects as indivisible, overlooking the components that constitute them. Understanding these components and their associated affordances provides valuable insights into an object's functionality, which is fundamental for performing a wide range of tasks. In this work, we introduce a novel real-world benchmark, InstructPart, comprising hand-labeled part segmentation annotations and task-oriented instructions to evaluate the performance of current models in understanding and executing part-level tasks within everyday contexts. Through our experiments, we demonstrate that task-oriented part segmentation remains a challenging problem, even for state-of-the-art Vision-Language Models (VLMs). In addition to our benchmark, we introduce a simple baseline that achieves a twofold performance improvement through fine-tuning with our dataset. With our dataset and benchmark, we aim to facilitate research on task-oriented part segmentation and enhance the applicability of VLMs across various domains, including robotics, virtual reality, information retrieval, and other related fields. Project website: https://zifuwan.github.io/InstructPart/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext df3c6625-a204-4ce7-8989-3a802fb4a4b6Cited by top-tier papers7
- SAM3-I: Segment Anything with InstructionsJingjing Li, Yue Feng, Yuchen Guo, Jincai Huang et al.ACL 2026 · 7 citations
- MonoFusion: Sparse-View 4D Reconstruction via Monocular FusionZihan Wang, Jeff Tan, Tarasha Khurana, Neehar Peri et al.ICCV 2025 · 5 citations
- ONLY: One-Layer Intervention Sufficiently Mitigates Hallucinations in Large Vision-Language ModelsZifu Wan, Ce Zhang, Silong Yong, Martin Q. Ma et al.ICCV 2025 · 2 citations
- Contact-guided Real2Sim from Monocular Video with Planar Scene PrimitivesZihan Wang, Jiashun Wang, Jeff Tan, Yiwen Zhao et al.ICLR 2026
- Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM DecodingLiu Yu, Can Chen, PING KUANG, Zhikun Feng et al.ICML 2026
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language ModelChunlin Yu, Hanqing Wang, Ye Shi, Haoyang Luo et al.CVPR 2025
- Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture AssemblyAditya Chetan, Eric Cai, Peeyush Kushwaha, Bharath Raj Nagoor Kani et al.CVPR 2026
- DetGPT: Detect What You Need via ReasoningRenjie Pi, Jiahui Gao, Shizhe Diao, Rui Pan et al.EMNLP 2023 · 57 citations
- VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language ModelsMingjie Xu, Jinpeng Chen, Yuzhi Zhao, Jason Chun Lok Li et al.AAAI 2026
- LASO: Language-Guided Affordance Segmentation on 3D ObjectYicong Li, Na Zhao, Junbin Xiao, Chun Feng et al.CVPR 2024
