HyPoGen: Optimization-Biased Hypernetworks for Generalizable Policy Generation
Hanxiang Ren, Li Sun, Xulong Wang, Pei Zhou, Zewen Wu, Siyan Dong, Difan Zou, Youyi Zheng, Yanchao Yang
Abstract
Policy learning through behavior cloning poses significant challenges, particularly when demonstration data is limited. In this work, we present HyPoGen, a novel optimization-biased hypernetwork for policy generation. The proposed hypernetwork learns to synthesize optimal policy parameters solely from task specifications -without accessing training data -by modeling policy generation as an approximation of the optimization process executed over a finite number of steps and assuming these specifications serve as a sufficient representation of the demonstration data. By incorporating structural designs that bias the hypernetwork towards optimization, we can improve its generalization capability while only training on source task demonstrations. During the feed-forward prediction pass, the hypernetwork effectively performs an optimization in the latent (compressed) policy space, which is then decoded into policy parameters for action prediction. Experimental results on locomotion and manipulation benchmarks show that HyPoGen significantly outperforms state-of-the-art methods in generating policies for unseen target tasks without any demonstrations, achieving higher success rates and underscoring the potential of optimization-biased hypernetworks in advancing generalizable policy generation. Our code and data are available at: https://github.com/ReNginx/HyPoGen .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Hyper-GoalNet: Goal-Conditioned Manipulation Policy Learning with HyperNetworksPei Zhou, Wanting Yao, Qian Luo, Xunzhe Zhou et al.NeurIPS 2025 · 4 citations
- Learning Diffusion Policy from Primitive Skills for Robot ManipulationZhihao Gu, Ming Yang, Difan Zou, Dong XuAAAI 2026
Builds on12
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Continual learning with hypernetworksJohannes von Oswald, Christian Henning, João Sacramento, Benjamin F. GreweICLR 2020 · 412 citations
- Learning the Pareto Front with HypernetworksAviv Navon, Aviv Shamsian, Ethan Fetaya, Gal ChechikICLR 2021 · 189 citations
- Observational Overfitting in Reinforcement LearningXingyou Song, Yiding Jiang, Stephen Tu, Yilun Du et al.ICLR 2020 · 148 citations
- Early Stopping in Deep Networks: Double Descent and How to Eliminate itReinhard Heckel, Fatih Furkan YilmazICLR 2021 · 55 citations
Related papers
- Make-An-Agent: A Generalizable Policy Network Generator with Behavior-Prompted DiffusionYongyuan Liang, Tingqiang Xu, Kaizhe Hu, Guangqi Jiang et al.NeurIPS 2024 · 15 citations
- Foundation Policies with Hilbert RepresentationsSeohong Park, Tobias Kreiman, Sergey LevineICML 2024 · 72 citations
- Generalizable Domain Adaptation for Sim-and-Real Policy Co-TrainingShuo Cheng, Liqian Ma, Zhenyang Chen, Ajay Mandlekar et al.NeurIPS 2025 · 16 citations
- Learning to Act from Actionless Videos through Dense CorrespondencesPo-Chen Ko, Jiayuan Mao, Yilun Du, Shao-Hua Sun et al.ICLR 2024 · 181 citations
- MoMaGen: Generating Demonstrations under Soft and Hard Constraints for Multi-Step Bimanual Mobile ManipulationChengshu Li, Mengdi Xu, Arpit Bahety, Hang Yin et al.ICLR 2026 · 17 citations
