FLIP: Flow-Centric Generative Planning as General-Purpose Manipulation World Model
Chongkai Gao, Haozhuo Zhang, Zhixuan Xu, Zhehao Cai, Lin Shao
Abstract
We aim to develop a model-based planning framework for world models that can be scaled with increasing model and data budgets for general-purpose manipulation tasks with only language and vision inputs. To this end, we present FLow-CentrIc generative Planning (FLIP), a model-based planning algorithm on visual space that features three key modules: 1) a multi-modal flow generation model as the general-purpose action proposal module; 2) a flow-conditioned video generation model as the dynamics module; and 3) a vision-language representation learning model as the value module. Given an initial image and language instruction as the goal, FLIP can progressively search for long-horizon flow and video plans that maximize the discounted return to accomplish the task. FLIP is able to synthesize long-horizon plans across objects, robots, and tasks with image flows as the general action representation, and the dense flow information also provides rich guidance for long-horizon video generation. In addition, the synthesized flow and video plans can guide the training of low-level control policies for robot execution. Experiments on diverse benchmarks demonstrate that FLIP can improve both the success rates and quality of long-horizon video plan synthesis and has the interactive world model property, opening up wider applications for future works. Video demos are on our website: https://nus-lins-lab.github.io/flipweb/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2b03f2c-1046-40bb-81e9-45ea86e6ccadCited by top-tier papers16
- Ctrl-World: A Controllable Generative World Model for Robot ManipulationYanjiang Guo, Lucy Xiaoyang Shi, Jianyu Chen, Chelsea FinnICLR 2026 · 163 citations
- Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching DistillationYunhong Lu, Yanhong Zeng, Haobo Li, Hao Ouyang et al.CVPR 2026 · 77 citations
- VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action ModelsChongkai Gao, Zixuan Liu, Zhenghao Chi, Junshan Huang et al.NeurIPS 2025 · 41 citations
- TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment VideosSeungjae Lee, Yoonkyo Jung, Inkook Chun, Yao-Chih Lee et al.CVPR 2026 · 17 citations
- StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State RepresentationMingyu Liu, Jiuhe Shu, Hui Chen, Zeju Li et al.CVPR 2026 · 14 citations
Builds on25
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
Related papers
- Video Language PlanningYilun Du, Sherry Yang, Pete Florence, Fei Xia et al.ICLR 2024 · 161 citations
- EVLP: Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-TuningXinyan Cai, Qiang Guan, Shiguang Wu, Dafeng Chi et al.ICLR 2026
- Compositional Foundation Models for Hierarchical PlanningAnurag Ajay, Seungwook Han, Yilun Du, Shuang Li et al.NeurIPS 2023 · 137 citations
- EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric FlowYixiang Chen, Peiyan Li, Yan Huang, Jiabing Yang et al.ICCV 2025 · 2 citations
- InstructFlow: Adaptive Symbolic Constraint-Guided Code Generation for Long-Horizon PlanningHaotian Chi, Zeyu Feng, Yueming Lyu, Chengqi Zheng et al.NeurIPS 2025 · 6 citations
