Beyond Object Recognition: A New Benchmark towards Object Concept Learning
Yonglu Li, Yue Xu, Xinyu Xu, Xiaohan Mao, Yuan Yao, Siqi Liu, Cewu Lu
Abstract
Understanding objects is a central building block of AI, especially for embodied AI. Even though object recognition excels with deep learning, current machines struggle to learn higher-level knowledge, e.g., what attributes an object has, and what we can do with it. Here, we propose a challenging Object Concept Learning (OCL) task to push the envelope of object understanding. It requires machines to reason out affordances and simultaneously give the reason: what attributes make an object possess these affordances. To support OCL, we build a densely annotated knowledge base including extensive annotations for three levels of object concept (category, attribute, affordance), and the clear causal relations of three levels. By analyzing the causal structure of OCL, we present a baseline, Object Concept Reasoning Network (OCRN). It leverages concept instantiation and causal intervention to infer the three levels. In experiments, OCRN effectively infers the object knowledge while following the causalities well. Our data and code are available at https://mvig-rhos.com/ocl.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 886db514-cf1c-4bd5-8e94-513b982c2090Cited by top-tier papers4
- M^3-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object SegmentationZixuan Chen, Jiaxin Li, Junxuan Liang, Liming Tan et al.CVPR 2025
- General Articulated Objects Manipulation in Real Images via Part-Aware Diffusion ProcessZhou Fang, Yong-Lu Li, Lixin Yang, Cewu LuNeurIPS 2024
- Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language ModelsNanxi Li, Xiang Wang, Yuanjie Chen, Haode Zhang et al.ICLR 2026
- The Labyrinth of Links: Navigating the Associative Maze of Multi-modal LLMsHong Li, Nanxi Li, Yuanjie Chen, Jianbin Zhu et al.ICLR 2025
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Long-Tailed Classification by Keeping the Good and Removing the Bad Momentum Causal EffectKaihua Tang, Jianqiang Huang, Hanwang ZhangNeurIPS 2020 · 533 citations
Related papers
- CoTDet: Affordance Knowledge Prompting for Task Driven Object DetectionJiajin Tang, Ge Zheng, Jingyi Yu, Sibei YangICCV 2023 · 47 citations
- 3D AffordanceNet: A Benchmark for Visual Object Affordance UnderstandingShengheng Deng, Xun Xu, Chaozheng Wu, Ke Chen et al.CVPR 2021
- Discovering Conceptual Knowledge with Analytic Ontology Templates for Articulated ObjectsJianhua Sun, Yuxuan Li, Longfei Xu, Jiude Wei et al.AAAI 2025 · 3 citations
- Grounding 3D Object Affordance with Language Instructions, Visual Observations and InteractionsHe Zhu, Quyu Kong, Kechun Xu, Xunlong Xia et al.CVPR 2025
- Affordance Transfer Learning for Human-Object Interaction DetectionZhi Hou, Baosheng Yu, Yu Qiao, Xiaojiang Peng et al.CVPR 2021
