Beyond Object Recognition: A New Benchmark towards Object Concept Learning
Yonglu Li, Yue Xu, Xinyu Xu, Xiaohan Mao, Yuan Yao, Siqi Liu, Cewu Lu
摘要
Understanding objects is a central building block of AI, especially for embodied AI. Even though object recognition excels with deep learning, current machines struggle to learn higher-level knowledge, e.g., what attributes an object has, and what we can do with it. Here, we propose a challenging Object Concept Learning (OCL) task to push the envelope of object understanding. It requires machines to reason out affordances and simultaneously give the reason: what attributes make an object possess these affordances. To support OCL, we build a densely annotated knowledge base including extensive annotations for three levels of object concept (category, attribute, affordance), and the clear causal relations of three levels. By analyzing the causal structure of OCL, we present a baseline, Object Concept Reasoning Network (OCRN). It leverages concept instantiation and causal intervention to infer the three levels. In experiments, OCRN effectively infers the object knowledge while following the causalities well. Our data and code are available at https://mvig-rhos.com/ocl.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- M^3-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object SegmentationZixuan Chen, Jiaxin Li, Junxuan Liang, Liming Tan 等CVPR 2025
- General Articulated Objects Manipulation in Real Images via Part-Aware Diffusion ProcessZhou Fang, Yong-Lu Li, Lixin Yang, Cewu LuNeurIPS 2024
- Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language ModelsNanxi Li, Xiang Wang, Yuanjie Chen, Haode Zhang 等ICLR 2026
- The Labyrinth of Links: Navigating the Associative Maze of Multi-modal LLMsHong Li, Nanxi Li, Yuanjie Chen, Jianbin Zhu 等ICLR 2025
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Long-Tailed Classification by Keeping the Good and Removing the Bad Momentum Causal EffectKaihua Tang, Jianqiang Huang, Hanwang ZhangNeurIPS 2020 · 被引用 533 次
相关 Paper
- CoTDet: Affordance Knowledge Prompting for Task Driven Object DetectionJiajin Tang, Ge Zheng, Jingyi Yu, Sibei YangICCV 2023 · 被引用 47 次
- 3D AffordanceNet: A Benchmark for Visual Object Affordance UnderstandingShengheng Deng, Xun Xu, Chaozheng Wu, Ke Chen 等CVPR 2021
- Discovering Conceptual Knowledge with Analytic Ontology Templates for Articulated ObjectsJianhua Sun, Yuxuan Li, Longfei Xu, Jiude Wei 等AAAI 2025 · 被引用 3 次
- Grounding 3D Object Affordance with Language Instructions, Visual Observations and InteractionsHe Zhu, Quyu Kong, Kechun Xu, Xunlong Xia 等CVPR 2025
- Affordance Transfer Learning for Human-Object Interaction DetectionZhi Hou, Baosheng Yu, Yu Qiao, Xiaojiang Peng 等CVPR 2021
