Lune

NeurIPS2025Top-tier venue

Scaling Up Parameter Generation: A Recurrent Diffusion Approach

Kai Wang, Dongwen Tang, Wangbo Zhao, Konstantin Schürholt, Zhangyang (Atlas) Wang, Yang You

2025Year
9Citations
5Top-tier citations

Abstract

Parameter generation has long struggled to match the scale of today’s large vision and language models, curbing its broader utility. In this paper, we introduce R ecurrent Diffusion for Large-Scale P arameter G eneration ( RPG ) , a novel framework that generates full neural network parameters—up to hundreds of millions —on a single GPU . Our approach first partitions a network’s parameters into non-overlapping ‘tokens’, each corresponding to a distinct portion of the model. A recurrent mechanism then learns the inter-token relationships, producing ‘prototypes’ which serve as conditions for a diffusion process that ultimately synthesizes the parameters. Across a spectrum of architectures and tasks—including ResNets, ConvNeXts and ViTs on ImageNet-1K and COCO, and even LoRA-based LLMs—RPG achieves performance on par with fully trained networks while avoiding excessive memory overhead. Notably, it generalizes beyond its training set to generate valid parameters for previously unseen tasks, highlighting its flexibility in open-ended scenarios. By overcoming the longstanding memory and scalability barriers, RPG serves as a critical advance in ‘ AI generating AI ’, potentially enabling efficient weight generation at scales previously deemed infeasible.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 87778c75-8697-4878-9c39-fa8afbda2f5a

Cited by top-tier papers5

Ask how each one uses it

Builds on31

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines