Building LLMs Like LEGO: Two-dimensional Architecture Reassembly of Large Language Models
Xingyu Wu, Yu Zhou, Kay Chen Tan
摘要
Pretrained large language models (LLMs) are typically reused as indivisible artifacts, adapted, merged, or ensembled as a whole. In this study, we show that LLMs can instead be structurally recomposed as modular building blocks to create new architectures without access to original training data. We introduce architecturelevel reassembly as a new reuse paradigm, in which Transformer blocks from heterogeneous models are treated as reusable components. This idea is formalized through a twodimensional reassembly space that supports both vertical recombination across depth and horizontal composition within layers. To make this space tractable, we propose a chromosomebased architectural encoding and perform a bilevel multi-objective evolutionary optimization over vertical structure and horizontal composition. To resolve representation incompatibility across heterogeneous blocks, we introduce lightweight glue layers trained via data-free knowledge distillation, enabling valid information flow without modifying pretrained parameters. Our results demonstrate that architecturelevel reassembly unlocks a new dimension of flexibility in model reuse, pointing toward a modular and evolutionary view of LLM design.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel 等NeurIPS 2023 · 被引用 999 次
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 被引用 741 次
- Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free LunchLe Yu, Bowen Yu, Haiyang Yu, Fei Huang 等ICML 2024 · 被引用 605 次
相关 Paper
- LS-Merge: Merging Language Models in Latent SpaceBedionita Soro, Aoxuan Silvia Zhang, Bruno Andreis, Jaehyeong Jo 等ICLR 2026
- Deep Model ReassemblyXingyi Yang, Daquan Zhou, Songhua Liu, Jingwen Ye 等NeurIPS 2022 · 被引用 162 次
- XPERT: Expert Knowledge Transfer for Effective Training of Language ModelsChang Liu, boyu shi, Xu Yang, Xin GengICML 2026 · 被引用 2 次
- Few-shot Task-agnostic Neural Architecture Search for Distilling Large Language ModelsDongkuan Xu, Subhabrata Mukherjee, Xiaodong Liu, Debadeepta Dey 等NeurIPS 2022 · 被引用 21 次
- Model LEGO: Creating Models Like Disassembling and Assembling Building BlocksJiacong Hu, Jing Gao, Jingwen Ye, Yang Gao 等NeurIPS 2024 · 被引用 1 次
