NNiT: Width-Agnostic Neural Network Generation with Structurally Aligned Weight Spaces
Jiwoo Kim, Swarajh Mehta, Hao-Lun Hsu, Hyunwoo Ryu, Yudong Liu, Miroslav Pajic
摘要
Generative modeling of neural network parameters is often tied to architectures because standard parameter representations rely on known weight-matrix dimensions. Generation is further complicated by permutation symmetries that allow networks to model similar input-output functions while having widely different, unaligned parameterizations. In this work, we introduce Neural Network Diffusion Transformers (NNiTs), which generate weights in a width-agnostic manner by tokenizing weight matrices into patches and modeling them as locally structured fields. We establish that Graph HyperNetworks (GHNs) with a convolutional neural network (CNN) decoder structurally align the weight space, creating the local correlation necessary for patch-based processing. Focusing on Multilayer Perceptrons (MLPs), where permutation symmetry is especially apparent, NNiTs generate fully functional networks across a range of architectures. Our approach jointly models discrete architecture tokens and continuous weight patches within a single sequence model. On ManiSkill3 robotics tasks, NNiT achieves success on architecture topologies unseen during training, while baseline approaches fail to generalize; the same pipeline also generalizes to MNIST classification beyond the robotic control setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper33
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Masked Diffusion Transformer is a Strong Image SynthesizerShanghua Gao, Pan Zhou, Ming-Ming Cheng, Shuicheng YanICCV 2023 · 被引用 290 次
- Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion TransformerShuang Wu, Youtian Lin, Yifei Zeng, Feihu Zhang 等NeurIPS 2024 · 被引用 251 次
相关 Paper
- Graph Neural Networks for Learning Equivariant Representations of Neural NetworksMiltiadis Kofinas, Boris Knyazev, Yan Zhang, Yunlu Chen 等ICLR 2024 · 被引用 57 次
- Text2Weight: Bridging Natural Language and Neural Network Weight SpacesBowen Tian, Wenshuo Chen, Zexi Li, Songning Lai 等ACM MM 2025
- DeepWeightFlow: Re-Basined Flow Matching for Generating Neural Network WeightsSaumya Gupta, Scott Biggs, Moritz Laber, Zohair Shafi 等ICLR 2026 · 被引用 5 次
- Graph Metanetworks for Processing Diverse Neural ArchitecturesDerek Lim, Haggai Maron, Marc T. Law, Jonathan Lorraine 等ICLR 2024 · 被引用 47 次
- Hyper-Transforming Latent Diffusion ModelsIgnacio Peis, Batuhan Koyuncu, Isabel Valera, Jes FrellsenICML 2025
