Seeing the Unseen: Visual Common Sense for Semantic Placement
Ram Ramrakhya, Aniruddha Kembhavi, Dhruv Batra, Zsolt Kira, Kuo-Hao Zeng, Luca Weihs
摘要
Computer vision tasks typically involve describing what is present in an image (e.g. classification, detection, segmentation, and captioning). We study a visual common sense task that requires understanding ‘what is not present’. Specif-ically, given an image (e.g. of a living room) and a name of an object (“cushion ”), a vision system is asked to predict semantically-meaningful regions (masks or bounding boxes) in the image where that object could be placed or is likely be placed by humans (e.g. on the sofa). We call this task: Se-mantic Placement (SP) and believe that such common-sense visual understanding is critical for assitive robots (tidying a house), AR devices (automatically rendering an object in the user's space), and visually-grounded chatbots with common sense. Studying the invisible is hard. Datasets for image description are typically constructed by curating relevant images (e.g. via image search with object names) and asking humans to annotate the contents of the image; neither of those two steps are straightforward for objects not present in the image. We overcome this challenge by operating in the opposite direction: we start with an image of an object in context, which is easy to find online, and then remove that ob-ject from the image via inpainting. This automated pipeline converts unstructured web data into a dataset comprising pairs of images with/without the object. With this proposed data generation pipeline, we collect a novel dataset, containing 1.3M images across 9 object categories. We then train a SP prediction model, called CLIP-UNet, on our dataset. The CLIP-UNet outperforms existing VLMs and baselines that combine semantic priors with object detectors, gener-alizes well to real-world and simulated images and exhibits semantics-aware reasoning for object placement. In our user studies, we find that the SP masks predicted by CLIP-UNet are favored 43.7% and 31.3% times when comparing against the 4 SP baselines on real and simulated images. In addition, leveraging SP mask predictions from CLIP-UNet enables downstream applications like building tidying robots in indoor environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Placeit3d: Language-Guided Object Placement in Real 3D ScenesAhmed Abdelreheem, Filippo Aleotti, Jamie Watson, Zawar Qureshi 等ICCV 2025 · 被引用 11 次
- From Semantics, Scene to Instance-awareness: Distilling Foundation Model for Open-vocabulary Grounded Situation RecognitionChen Cai, Tianyi Liu, Jianjun Gao, Wenyang Liu 等ACM MM 2025 · 被引用 2 次
- Joint Navigation and Manipulation Planning with 3D Interaction ChainsKeming Zhang, Sixian Zhang, Xinhang Song, Hongyu Wang 等ICML 2026
它引用的顶会 Paper29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- Amodal Panoptic SegmentationRohit Mohan, Abhinav ValadaCVPR 2022 · 被引用 49 次
- A-CAP: Anticipation Captioning with Commonsense KnowledgeDuc Minh Vo, Quoc-An Luong, Akihiro Sugimoto, Hideki NakayamaCVPR 2023
- Zero-Shot Composed Image Retrieval with Textual InversionAlberto Baldrati, Lorenzo Agnolucci, Marco Bertini, Alberto Del BimboICCV 2023 · 被引用 214 次
- 3D AffordanceNet: A Benchmark for Visual Object Affordance UnderstandingShengheng Deng, Xun Xu, Chaozheng Wu, Ke Chen 等CVPR 2021
- Towards Open-Vocabulary Semantic Segmentation Without Semantic LabelsHeeseong Shin, Chaehyun Kim, Sunghwan Hong, Seokju Cho 等NeurIPS 2024 · 被引用 32 次
