Understanding 3D Object Interaction from a Single Image
Shengyi Qian, David F. Fouhey
摘要
Humans can easily understand a single image as depicting multiple potential objects permitting interaction. We use this skill to plan our interactions with the world and accelerate understanding new objects without engaging in interaction. In this paper, we would like to endow machines with the similar ability, so that intelligent agents can better explore the 3D scene or manipulate objects. Our approach is a transformer-based model that predicts the 3D location, physical properties and affordance of objects. To power this model, we collect a dataset with Internet videos, egocentric videos and indoor images to train and validate our approach. Our model yields strong performance on our data, and generalizes well to robotics data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- DIPO: Dual-State Images Controlled Articulated Object Generation Powered by Diverse DataRuiqi Wu, Xinjie Wang, Liu Liu, Chun-Le Guo 等NeurIPS 2025 · 被引用 21 次
- HoloScene: Simulation-Ready Interactive 3D Worlds from a Single VideoHongchi Xia, Chih-Hao Lin, Hao-Yu Hsu, Quentin Leboutet 等NeurIPS 2025 · 被引用 18 次
- Dexterous World ModelsByungjun Kim, Taeksoo Kim, Junyoung Lee, Hanbyul JooCVPR 2026 · 被引用 17 次
- RAGNet: Large-Scale Reasoning-Based Affordance Segmentation Benchmark Towards General GraspingDongming Wu, Yanping Fu, Saike Huang, Yingfei Liu 等ICCV 2025 · 被引用 2 次
- Beyond the Frame: Generating 360° Panoramic Videos from Perspective VideosRundong Luo, Matthew Wallingford, Ali Farhadi, Noah Snavely 等ICCV 2025 · 被引用 2 次
它引用的顶会 Paper25
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil 等NeurIPS 2020 · 被引用 4,036 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- Omnidata: A Scalable Pipeline for Making Multi-Task Mid-Level Vision Datasets from 3D ScansAinaz Eftekhar, Alexander Sax, Jitendra Malik, Amir ZamirICCV 2021 · 被引用 422 次
相关 Paper
- Grounding 3D Object Affordance from 2D Interactions in ImagesYuhang Yang, Wei Zhai, Hongchen Luo, Yang Cao 等ICCV 2023 · 被引用 69 次
- Grounding 3D Object Affordance with Language Instructions, Visual Observations and InteractionsHe Zhu, Quyu Kong, Kechun Xu, Xunlong Xia 等CVPR 2025
- GanHand: Predicting Human Grasp Affordances in Multi-Object ScenesEnric Corona, Albert Pumarola, Guillem Alenyà, Francesc Moreno-Noguer 等CVPR 2020
- TB-HSU: Hierarchical 3D Scene Understanding with Contextual AffordancesWenting Xu, Viorela Ila, Luping Zhou, Craig T. JinAAAI 2025 · 被引用 4 次
- Neural-Logic Human-Object Interaction DetectionLiulei Li, Jianan Wei, Wenguan Wang, Yi YangNeurIPS 2023 · 被引用 54 次
