H2OFlow: Grounding Human-Object Affordances with 3D Generative Models and Dense Diffused Flows
Harry Zhang, Luca Carlone
Abstract
Understanding how humans interact with the surrounding environment, and specifically reasoning about object interactions and affordances, is a critical challenge in computer vision, robotics, and AI. Current approaches often depend on labor-intensive, hand-labeled datasets capturing real-world or simulated human-object interaction (HOI) tasks, which are costly and time-consuming to produce. Furthermore, most existing methods for 3D affordance understanding are limited to contact-based analysis, neglecting other essential aspects of human-object interactions, such as orientation (, humans might have a preferential orientation with respect certain objects, such as a TV) and spatial occupancy (, humans are more likely to occupy certain regions around an object, like the front of a microwave rather than its back). To address these limitations, we introduce H2OFlow, a novel framework that comprehensively learns 3D HOI affordances -- encompassing contact, orientation, and spatial occupancy -- using only synthetic data generated from 3D generative models. H2OFlow employs a dense 3D-flow-based representation, learned through a dense diffusion process operating on point clouds. This learned flow enables the discovery of rich 3D affordances without the need for human annotations. Through extensive quantitative and qualitative evaluations, we demonstrate that H2OFlow generalizes effectively to real-world objects and surpasses prior methods that rely on manual annotations or mesh-based representations in modeling 3D affordance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext abb7b718-02b2-4182-b2af-733d0c4c50f6Builds on36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- DAViD: Modeling Dynamic Affordance of 3D Objects Using Pre-Trained Video Diffusion ModelsHyeonwoo Kim, Sangwon Baik, Hanbyul JooICCV 2025 · 1 citation
- G3Flow: Generative 3D Semantic Flow for Pose-aware and Generalizable Object ManipulationTianxing Chen, Yao Mu, Zhixuan Liang, Zanxin Chen et al.CVPR 2025
- CHORUS: Learning Canonicalized 3D Human-Object Spatial Relations from Unbounded Synthesized ImagesSookwan Han, Hanbyul JooICCV 2023 · 19 citations
- NIFTY: Neural Object Interaction Fields for Guided Human Motion SynthesisNilesh Kulkarni, Davis Rempe, Kyle Genova, Abhijit Kundu et al.CVPR 2024
- Visual Relation Diffusion for Human-Object Interaction DetectionPing Cao, Yepeng Tang, Chunjie Zhang, Xiaolong Zheng et al.ICCV 2025 · 1 citation
