PourIt!: Weakly-supervised Liquid Perception from a Single Image for Visual Closed-Loop Robotic Pouring
Haitao Lin, Yanwei Fu, Xiangyang Xue
摘要
Liquid perception is critical for robotic pouring tasks. It usually requires the robust visual detection of flowing liquid. However, while recent works have shown promising results in liquid perception, they typically require labeled data for model training, a process that is both time-consuming and reliant on human labor. To this end, this paper proposes a simple yet effective framework PourIt!, to serve as a tool for robotic pouring tasks. We design a simple data collection pipeline that only needs image-level labels to reduce the reliance on tedious pixel-wise annotations. Then, a binary classification model is trained to generate Class Activation Map (CAM) that focuses on the visual difference between these two kinds of collected data, i.e., the existence of liquid drop or not. We also devise a feature contrast strategy to improve the quality of the CAM, thus entirely and tightly covering the actual liquid regions. Then, the container pose is further utilized to facilitate the 3D point cloud recovery of the detected liquid region. Finally, the liquid-to-container distance is calculated for visual closed-loop control of the physical robot. To validate the effectiveness of our proposed method, we also contribute a novel dataset for our task and name it PourIt! dataset. Extensive results on this dataset and physical Franka robot have shown the utility and effectiveness of our method in the robotic pouring tasks. Our dataset, code and pre-trained models will be available on the project page 1 . Input: A single RGB-D image 2D liquid perception & 6-DoF object pose 6-DoF object pose & 3D object size Open-loop with only bottle-to-bottle distance Closed-loop with Liquid-to-bottle distance Robotics Manipulation Ours Prev. method 3D liquid modeling 3D point cloud Missing depth Refraction of light
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Towards Long-Horizon Vision-Language-Action System: Reasoning, Acting and MemoryDaixun Li, Yusi Zhang, Mingxiang Cao, Donglai Liu 等ICCV 2025 · 被引用 2 次
- A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual GroundingZhenyang Liu, Sixiao Zheng, Siyu Chen, Cairong Zhao 等ACM MM 2025 · 被引用 1 次
- Phys-Liquid: A Physics-Informed Dataset for Estimating 3D Geometry and Volume of Transparent Deformable LiquidsKe Ma, Yizhou Fang, Jean-Baptiste Weibel, Shuai Tan 等AAAI 2026
- CAP-Net: A Unified Network for 6D Pose and Size Estimation of Categorical Articulated Parts from a Single RGB-D ImageJingshun Huang, Haitao Lin, Tianyu Wang, Yanwei Fu 等CVPR 2025
它引用的顶会 Paper12
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- Learning Affinity from Attention: End-to-End Weakly-Supervised Semantic Segmentation with TransformersLixiang Ru, Yibing Zhan, Baosheng Yu, Bo DuCVPR 2022 · 被引用 257 次
- Integral Object Mining via Online Attention AccumulationPeng-Tao Jiang, Qibin Hou, Yang Cao, Ming-Ming Cheng 等ICCV 2019 · 被引用 246 次
- Class Re-Activation Maps for Weakly-Supervised Semantic SegmentationZhaozheng Chen, Tan Wang, Xiongwei Wu, Xian-Sheng Hua 等CVPR 2022 · 被引用 223 次
相关 Paper
- Image Based Reconstruction of Liquids from 2D Surface DetectionsFlorian Richter, Ryan K. Orosco, Michael C. YipCVPR 2022 · 被引用 7 次
- The Functional Correspondence ProblemZihang Lai, Senthil Purushwalkam, Abhinav GuptaICCV 2021 · 被引用 23 次
- YouRefIt: Embodied Reference Understanding with Language and GestureYixin Chen, Qing Li, Deqian Kong, Yik Lun Kei 等ICCV 2021 · 被引用 57 次
- LiquImager: Fine-grained Liquid Identification and Container Imaging System with COTS WiFi DevicesFei Shang, Panlong Yang, Dawei Yan, Sijia Zhang 等UbiComp 2024 · 被引用 17 次
- Fixing Malfunctional Objects With Learned Physical Simulation and Functional PredictionYining Hong, Kaichun Mo, Li Yi, Leonidas J. Guibas 等CVPR 2022 · 被引用 4 次
