Scaling Up Parameter Generation: A Recurrent Diffusion Approach
Kai Wang, Dongwen Tang, Wangbo Zhao, Konstantin Schürholt, Zhangyang (Atlas) Wang, Yang You
Abstract
Parameter generation has long struggled to match the scale of today’s large vision and language models, curbing its broader utility. In this paper, we introduce R ecurrent Diffusion for Large-Scale P arameter G eneration ( RPG ) , a novel framework that generates full neural network parameters—up to hundreds of millions —on a single GPU . Our approach first partitions a network’s parameters into non-overlapping ‘tokens’, each corresponding to a distinct portion of the model. A recurrent mechanism then learns the inter-token relationships, producing ‘prototypes’ which serve as conditions for a diffusion process that ultimately synthesizes the parameters. Across a spectrum of architectures and tasks—including ResNets, ConvNeXts and ViTs on ImageNet-1K and COCO, and even LoRA-based LLMs—RPG achieves performance on par with fully trained networks while avoiding excessive memory overhead. Notably, it generalizes beyond its training set to generate valid parameters for previously unseen tasks, highlighting its flexibility in open-ended scenarios. By overcoming the longstanding memory and scalability barriers, RPG serves as a critical advance in ‘ AI generating AI ’, potentially enabling efficient weight generation at scales previously deemed infeasible.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 87778c75-8697-4878-9c39-fa8afbda2f5aCited by top-tier papers5
- GeoSANE: Learning Geospatial Representations from Models, Not DataJoëlle Hanna, Damian Falk, Stella X. Yu, Damian BorthCVPR 2026 · 1 citation
- WeightCLIP: Aligning Datasets and Models for Weight Space LearningAron Asefaw, Konstantinos Tzevelekakis, Damian Falk, Léo Meynent et al.ICML 2026
- LoFA: Learning to Predict Personalized Prior for Fast Adaptation of Visual Generative ModelsYiming Hao, Mutian Xu, Chongjie Ye, Jie Qin et al.CVPR 2026
- NNiT: Width-Agnostic Neural Network Generation with Structurally Aligned Weight SpacesJiwoo Kim, Swarajh Mehta, Hao-Lun Hsu, Hyunwoo Ryu et al.ICML 2026
- Parameter Manifold PurificationJiacong Hu, Jinxun Wu, Shengxuming Zhang, Shunyu Liu et al.ICML 2026
Builds on31
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- LoRAGen: Structure-Aware Weight Space Learning for LoRA GenerationHao Huang, Jingtao Ding, Mengqi Liao, Xin Wang et al.ICLR 2026
- Neural Residual Diffusion Models for Deep Scalable Vision GenerationZhiyuan Ma, Liangliang Zhao, Biqing Qi, Bowen ZhouNeurIPS 2024 · 15 citations
- Text Generation with Diffusion Language Models: A Pre-training Approach with Continuous Paragraph DenoiseZhenghao Lin, Yeyun Gong, Yelong Shen, Tong Wu et al.ICML 2023 · 107 citations
- Decentralized Diffusion ModelsDavid McAllister, Matthew Tancik, Jiaming Song, Angjoo KanazawaCVPR 2025
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski et al.ICML 2023 · 848 citations
