HyPoGen: Optimization-Biased Hypernetworks for Generalizable Policy Generation
Hanxiang Ren, Li Sun, Xulong Wang, Pei Zhou, Zewen Wu, Siyan Dong, Difan Zou, Youyi Zheng, Yanchao Yang
摘要
Policy learning through behavior cloning poses significant challenges, particularly when demonstration data is limited. In this work, we present HyPoGen, a novel optimization-biased hypernetwork for policy generation. The proposed hypernetwork learns to synthesize optimal policy parameters solely from task specifications -without accessing training data -by modeling policy generation as an approximation of the optimization process executed over a finite number of steps and assuming these specifications serve as a sufficient representation of the demonstration data. By incorporating structural designs that bias the hypernetwork towards optimization, we can improve its generalization capability while only training on source task demonstrations. During the feed-forward prediction pass, the hypernetwork effectively performs an optimization in the latent (compressed) policy space, which is then decoded into policy parameters for action prediction. Experimental results on locomotion and manipulation benchmarks show that HyPoGen significantly outperforms state-of-the-art methods in generating policies for unseen target tasks without any demonstrations, achieving higher success rates and underscoring the potential of optimization-biased hypernetworks in advancing generalizable policy generation. Our code and data are available at: https://github.com/ReNginx/HyPoGen .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Hyper-GoalNet: Goal-Conditioned Manipulation Policy Learning with HyperNetworksPei Zhou, Wanting Yao, Qian Luo, Xunzhe Zhou 等NeurIPS 2025 · 被引用 4 次
- Learning Diffusion Policy from Primitive Skills for Robot ManipulationZhihao Gu, Ming Yang, Difan Zou, Dong XuAAAI 2026
它引用的顶会 Paper12
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- Continual learning with hypernetworksJohannes von Oswald, Christian Henning, João Sacramento, Benjamin F. GreweICLR 2020 · 被引用 412 次
- Learning the Pareto Front with HypernetworksAviv Navon, Aviv Shamsian, Ethan Fetaya, Gal ChechikICLR 2021 · 被引用 189 次
- Observational Overfitting in Reinforcement LearningXingyou Song, Yiding Jiang, Stephen Tu, Yilun Du 等ICLR 2020 · 被引用 148 次
- Early Stopping in Deep Networks: Double Descent and How to Eliminate itReinhard Heckel, Fatih Furkan YilmazICLR 2021 · 被引用 55 次
相关 Paper
- Make-An-Agent: A Generalizable Policy Network Generator with Behavior-Prompted DiffusionYongyuan Liang, Tingqiang Xu, Kaizhe Hu, Guangqi Jiang 等NeurIPS 2024 · 被引用 15 次
- Foundation Policies with Hilbert RepresentationsSeohong Park, Tobias Kreiman, Sergey LevineICML 2024 · 被引用 72 次
- Generalizable Domain Adaptation for Sim-and-Real Policy Co-TrainingShuo Cheng, Liqian Ma, Zhenyang Chen, Ajay Mandlekar 等NeurIPS 2025 · 被引用 16 次
- Learning to Act from Actionless Videos through Dense CorrespondencesPo-Chen Ko, Jiayuan Mao, Yilun Du, Shao-Hua Sun 等ICLR 2024 · 被引用 181 次
- MoMaGen: Generating Demonstrations under Soft and Hard Constraints for Multi-Step Bimanual Mobile ManipulationChengshu Li, Mengdi Xu, Arpit Bahety, Hang Yin 等ICLR 2026 · 被引用 17 次
