Primary-Fine Decoupling for Action Generation in Robotic Imitation
Xiaohan Lei, Min Wang, Wengang Zhou, Xingyu Lu, Houqiang Li
摘要
Multi-modal distribution in robotic manipulation action sequences poses critical challenges for imitation learning. To this end, existing approaches often model the action space as either a discrete set of tokens or a continuous, latent-variable distribution. However, both approaches present trade-offs: some methods discretize actions into tokens and therefore lose fine-grained action variations, while others generate continuous actions in a single stage tend to produce unstable mode transitions. To address these limitations, we propose Primary-Fine Decoupling for Action Generation (PF-DAG), a two-stage framework that decouples coarse action consistency from fine-grained variations. First, we compress action chunks into a small set of discrete modes, enabling a lightweight policy to select consistent coarse modes and avoid mode bouncing. Second, a mode conditioned MeanFlow policy is learned to generate high-fidelity continuous actions. Theoretically, we prove PF-DAG’s two-stage design achieves a strictly lower MSE bound than single-stage generative policies. Empirically, PF-DAG outperforms state-of-the-art baselines across 56 tasks from Adroit, DexArt, and MetaWorld benchmarks. It further generalizes to real-world tactile dexterous manipulation tasks. Our work demonstrates that explicit mode-level decoupling enables both robust multi-modal modeling and reactive closed-loop control for robotic manipulation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- MSP: Probabilistically Consistent Multi-Scale Action GenerationZhixuan Lin, Gengqi Liu, Chao Zheng, Gao Lin 等ICML 2026
- DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile AdapterXukun Li, Yu Sun, Lei Zhang, Bo-Sheng Huang 等ICML 2026
- Behavioral Mode Discovery for Fine-tuning Multimodal Generative PoliciesAlberta Longhini, David Emukpere, Jean-Michel Renders, Seungsu KimICML 2026 · 被引用 1 次
- FreqPolicy: Frequency Autoregressive Visuomotor Policy with Continuous TokensYiming Zhong, Yumeng Liu, Chuyang Xiao, Zemin Yang 等NeurIPS 2025 · 被引用 16 次
- FoAM: Foresight-Augmented Multi-Task Imitation Policy for Robotic ManipulationLitao Liu, Wentao Wang, Yifan Han, Zhuoli Xie 等AAAI 2026 · 被引用 4 次
