H-GAR: A Hierarchical Interaction Framework via Goal-Driven Observation-Action Refinement for Robotic Manipulation
Yijie Zhu, Rui Shao, Ziyang Liu, Jie He, Jizhihui Liu, Jiuru Wang, Zitong Yu
Abstract
Unified video and action prediction models hold great potential for robotic manipulation, as future observations offer contextual cues for planning, while actions reveal how interactions shape the environment. However, most existing approaches treat observation and action generation in a monolithic and goal-agnostic manner, often leading to semantically misaligned predictions and incoherent behaviors. To this end, we propose H-GAR, a Hierarchical interaction framework via Goal-driven observation-Action Refinement. To anchor prediction to the task objective, H-GAR first produces a goal observation and a coarse action sketch that outline a high-level route toward the goal. To enable explicit interaction between observation and action under the guidance of the goal observation for more coherent decision-making, we devise two synergistic modules. (1) Goal-Conditioned Observation Synthesizer (GOS) synthesizes intermediate observations based on the coarse-grained actions and the predicted goal observation. (2) Interaction-Aware Action Refiner (IAAR) refines coarse actions into fine-grained, goal-consistent actions by leveraging feedback from the intermediate observations and a Historical Action Memory Bank that encodes prior actions to ensure temporal consistency. By integrating goal grounding with explicit action-observation interaction in a coarse-to-fine manner, H-GAR enables more accurate manipulation. Extensive experiments on both simulation and real-world robotic manipulation tasks demonstrate that H-GAR achieves state-of-the-art performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3f89ed37-59bf-419d-b284-7622dafc2d2aCited by top-tier papers5
- Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic ManipulationZaijing Li, Bing Hu, Rui Shao, Gongwei Chen et al.CVPR 2026 · 23 citations
- HiconAgent: History Context-aware Policy Optimization for GUI AgentsXurui Zhou, Gongwei Chen, Yuquan Xie, Zaijing Li et al.CVPR 2026 · 11 citations
- ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic ManipulationWei Li, Jizhihui Liu, Yixing Li, Junwen Tong et al.CVPR 2026 · 8 citations
- HATS: Hardness-Aware Trajectory Synthesis for GUI AgentsRui Shao, Ruize Gao, Bin Xie, Yixing Li et al.CVPR 2026 · 7 citations
- FLOW: Feature-Level Optimal Warping for Generalized Remote Physiological Measurementbo zhao, Junzhe Cao, Dan Guo, Dongmin Huang et al.CVPR 2026 · 2 citations
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng et al.NeurIPS 2024 · 758 citations
- Learning Universal Policies via Text-Guided Video GenerationYilun Du, Sherry Yang, Bo Dai, Hanjun Dai et al.NeurIPS 2023 · 742 citations
- MaskGIT: Masked Generative Image TransformerHuiwen Chang, Han Zhang, Lu Jiang, Ce Liu et al.CVPR 2022 · 346 citations
Related papers
- Hand-Centric Motion Refinement for 3D Hand-Object Interaction via Hierarchical Spatial-Temporal ModelingYuze Hao, Jianrong Zhang, Tao Zhuo, Fuan Wen et al.AAAI 2024 · 7 citations
- InterPrior: Scaling Generative Control for Physics-Based Human-Object InteractionsSirui Xu, Samuel Schulter, Morteza Ziyadi, Xialin He et al.CVPR 2026 · 14 citations
- Hierarchical Foresight: Self-Supervised Learning of Long-Horizon Tasks via Visual Subgoal GenerationSuraj Nair, Chelsea FinnICLR 2020 · 152 citations
- Hierarchical Video Prediction Using Relational Layouts for Human-Object InteractionsNavaneeth Bodla, Gaurav Shrivastava, Rama Chellappa, Abhinav ShrivastavaCVPR 2021
- HOSIG: Full-Body Human-Object-Scene Interaction Generation with Hierarchical Scene PerceptionWei Yao, Yunlian Sun, Hongwen Zhang, Yebin Liu et al.AAAI 2026 · 4 citations
