G3Flow: Generative 3D Semantic Flow for Pose-aware and Generalizable Object Manipulation
Tianxing Chen, Yao Mu, Zhixuan Liang, Zanxin Chen, Shijia Peng, Qiangyu Chen, Mingkun Xu, Ruizhen Hu, Hongyuan Zhang, Xuelong Li, Ping Luo
Abstract
Recent advances in imitation learning for 3D robotic manipulation have shown promising results with diffusionbased policies. However, achieving human-level dexterity requires seamless integration of geometric precision and semantic understanding. We present G3Flow, a novel framework that constructs real-time semantic flow, a dynamic, object-centric 3D semantic representation by leveraging foundation models. Our approach uniquely combines 3D generative models for digital twin creation, vision foundation models for semantic feature extraction, and robust pose tracking for continuous semantic flow updates. This integration enables complete semantic understanding even under occlusions while eliminating manual annotation requirements. By incorporating semantic flow into diffusion policies, extensive experiments across five simulation tasks show that G3Flow consistently outperforms existing approaches, achieving up to 68.3% and 50.1% success rates on terminal-constrained manipulation and crossobject generalization respectively. Our results demonstrate the effectiveness of G3Flow in enhancing real-time dynamic semantic feature understanding for robotic policies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 948da3e5-ab1d-493f-8a04-3abcdbca025eCited by top-tier papers11
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic ManipulationTianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai et al.ICML 2026 · 394 citations
- World Guidance: World Modeling in Condition Space for Action GenerationYue Su, Sijin Chen, Haixin Shi, Mingyu Liu et al.ICML 2026 · 26 citations
- RoboWheel: A Data Engine from Real-World Human Demonstrations for Cross-Embodiment Robotic LearningYuhong Zhang, Zihan Gao, Shengpeng Li, Ling-Hao Chen et al.CVPR 2026 · 11 citations
- Action-Geometry Prediction with 3D Geometric Prior for Bimanual ManipulationChongyang Xu, Haipeng Li, Shen Cheng, Haoqiang Fan et al.CVPR 2026 · 10 citations
- VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic RoutingYixiao Wang, Mingxiao Huo, Zhixuan Liang, Yushi Du et al.ICLR 2026 · 5 citations
Builds on11
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Planning with Diffusion for Flexible Behavior SynthesisMichael Janner, Yilun Du, Joshua B. Tenenbaum, Sergey LevineICML 2022 · 1,115 citations
- GET3D: A Generative Model of High Quality 3D Textured Shapes Learned from ImagesJun Gao, Tianchang Shen, Zian Wang, Wenzheng Chen et al.NeurIPS 2022 · 661 citations
- DreamFusion: Text-to-3D using 2D DiffusionBen Poole, Ajay Jain, Jonathan T. Barron, Ben MildenhallICLR 2023 · 463 citations
- FoundationPose: Unified 6D Pose Estimation and Tracking of Novel ObjectsBowen Wen, Wei Yang, Jan Kautz, Stan BirchfieldCVPR 2024 · 215 citations
Related papers
- H2OFlow: Grounding Human-Object Affordances with 3D Generative Models and Dense Diffused FlowsHarry Zhang, Luca CarloneICLR 2026 · 2 citations
- Efficient Hybrid SE(3)-Equivariant Visuomotor Flow Policy via Spherical Harmonics for Robot ManipulationQinglun Zhang, Shen Cheng, Tian Dan, Haoqiang Fan et al.CVPR 2026 · 1 citation
- GenFlowRL: Shaping Rewards with Generative Object-Centric Flow in Visual Reinforcement LearningKelin Yu, Sheng Zhang, Harshit Soora, Furong Huang et al.ICCV 2025
- EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric FlowYixiang Chen, Peiyan Li, Yan Huang, Jiabing Yang et al.ICCV 2025 · 2 citations
- GraspLDP: Towards Generalizable Grasping Policy via Latent DiffusionEnda Xiang, Haoxiang Ma, Xinzhu Ma, Zicheng Liu et al.CVPR 2026 · 2 citations
