PourIt!: Weakly-supervised Liquid Perception from a Single Image for Visual Closed-Loop Robotic Pouring
Haitao Lin, Yanwei Fu, Xiangyang Xue
Abstract
Liquid perception is critical for robotic pouring tasks. It usually requires the robust visual detection of flowing liquid. However, while recent works have shown promising results in liquid perception, they typically require labeled data for model training, a process that is both time-consuming and reliant on human labor. To this end, this paper proposes a simple yet effective framework PourIt!, to serve as a tool for robotic pouring tasks. We design a simple data collection pipeline that only needs image-level labels to reduce the reliance on tedious pixel-wise annotations. Then, a binary classification model is trained to generate Class Activation Map (CAM) that focuses on the visual difference between these two kinds of collected data, i.e., the existence of liquid drop or not. We also devise a feature contrast strategy to improve the quality of the CAM, thus entirely and tightly covering the actual liquid regions. Then, the container pose is further utilized to facilitate the 3D point cloud recovery of the detected liquid region. Finally, the liquid-to-container distance is calculated for visual closed-loop control of the physical robot. To validate the effectiveness of our proposed method, we also contribute a novel dataset for our task and name it PourIt! dataset. Extensive results on this dataset and physical Franka robot have shown the utility and effectiveness of our method in the robotic pouring tasks. Our dataset, code and pre-trained models will be available on the project page 1 . Input: A single RGB-D image 2D liquid perception & 6-DoF object pose 6-DoF object pose & 3D object size Open-loop with only bottle-to-bottle distance Closed-loop with Liquid-to-bottle distance Robotics Manipulation Ours Prev. method 3D liquid modeling 3D point cloud Missing depth Refraction of light
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f2514e27-7900-47aa-875a-753e00dd8d48Cited by top-tier papers4
- Towards Long-Horizon Vision-Language-Action System: Reasoning, Acting and MemoryDaixun Li, Yusi Zhang, Mingxiang Cao, Donglai Liu et al.ICCV 2025 · 2 citations
- A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual GroundingZhenyang Liu, Sixiao Zheng, Siyu Chen, Cairong Zhao et al.ACM MM 2025 · 1 citation
- Phys-Liquid: A Physics-Informed Dataset for Estimating 3D Geometry and Volume of Transparent Deformable LiquidsKe Ma, Yizhou Fang, Jean-Baptiste Weibel, Shuai Tan et al.AAAI 2026
- CAP-Net: A Unified Network for 6D Pose and Size Estimation of Categorical Articulated Parts from a Single RGB-D ImageJingshun Huang, Haitao Lin, Tianyu Wang, Yanwei Fu et al.CVPR 2025
Builds on12
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Learning Affinity from Attention: End-to-End Weakly-Supervised Semantic Segmentation with TransformersLixiang Ru, Yibing Zhan, Baosheng Yu, Bo DuCVPR 2022 · 257 citations
- Integral Object Mining via Online Attention AccumulationPeng-Tao Jiang, Qibin Hou, Yang Cao, Ming-Ming Cheng et al.ICCV 2019 · 246 citations
- Class Re-Activation Maps for Weakly-Supervised Semantic SegmentationZhaozheng Chen, Tan Wang, Xiongwei Wu, Xian-Sheng Hua et al.CVPR 2022 · 223 citations
Related papers
- Image Based Reconstruction of Liquids from 2D Surface DetectionsFlorian Richter, Ryan K. Orosco, Michael C. YipCVPR 2022 · 7 citations
- The Functional Correspondence ProblemZihang Lai, Senthil Purushwalkam, Abhinav GuptaICCV 2021 · 23 citations
- YouRefIt: Embodied Reference Understanding with Language and GestureYixin Chen, Qing Li, Deqian Kong, Yik Lun Kei et al.ICCV 2021 · 57 citations
- LiquImager: Fine-grained Liquid Identification and Container Imaging System with COTS WiFi DevicesFei Shang, Panlong Yang, Dawei Yan, Sijia Zhang et al.UbiComp 2024 · 17 citations
- Fixing Malfunctional Objects With Learned Physical Simulation and Functional PredictionYining Hong, Kaichun Mo, Li Yi, Leonidas J. Guibas et al.CVPR 2022 · 4 citations
