Independence Tests for Language Models
Sally Zhu, Ahmed M. Ahmed, Rohith Kuditipudi, Percy Liang
摘要
We consider the following problem: given the weights of two models, can we test whether they were trained independently-i.e., from independent random initializations? We consider two settings: constrained and unconstrained. In the constrained setting, we make assumptions about model architecture and training and propose a family of statistical tests that yield exact p-values with respect to the null hypothesis that the models are trained from independent random initializations. These p-values are valid regardless of the composition of either model's training data; we compute them by simulating exchangeable copies of each model under our assumptions and comparing various similarity measures of weights and activations between the original two models versus these copies. We report the p-values from these tests on pairs of 21 open-weight models (210 total pairs) and find we correctly identify all pairs of non-independent models. Notably, our tests remain effective even if one of the models was fine-tuned for many tokens. In the unconstrained setting, where we make no assumptions about training procedures, can change model architecture, and allow for adversarial evasion attacks, the previous tests no longer work. Instead, we propose a new test which matches hidden activations between two models, and use it to construct a test that is robust to adversarial transformations and to changes in model architecture. The test can also perform localized testing: identifying specific non-independent components of models. Though we no longer obtain exact p-values from this test, empirically we find it behaves as one and reliably distinguishes nonindependent models. Notably, we can use the test to identify specific parts of one model that are derived from another (e.g., how Llama 3.1-8B was pruned to initialize Llama 3.2-3B, or shared layers between Mistral-7B and StripedHyena-7B), and it is even robust to retraining individual layers of either model from scratch.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- LLM DNA: Tracing Model Evolution via Functional RepresentationsZhaomin Wu, Haodong Zhao, Ziyang Wang, Jizhou Guo 等ICLR 2026 · 被引用 21 次
- Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model OutputsYiwei Chen, Soumyadeep Pal, Yimeng Zhang, Qing Qu 等ICLR 2026 · 被引用 15 次
- Blackbox Model Provenance via Palimpsestic Membership InferenceRohith Kuditipudi, Jing Huang, Sally Zhu, Diyi Yang 等NeurIPS 2025 · 被引用 12 次
- ZipLLM: Efficient LLM Storage via Model-Aware Synergistic Data Deduplication and CompressionZirui Wang, Tingfeng Lan, Zhaoyuan Su, Juncheng Yang 等NSDI 2026 · 被引用 8 次
- Adaptive Multiscale Binary Expansion Tests for IndependenceYang Yang, Duo Zheng, Sandeep Jain, Kai Zhang 等ICML 2026
它引用的顶会 Paper10
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin 等NeurIPS 2023 · 被引用 1,975 次
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz 等ICML 2023 · 被引用 854 次
- Sheared LLaMA: Accelerating Language Model Pre-training via Structured PruningMengzhou Xia, Tianyu Gao, Zhiyuan Zeng, Danqi ChenICLR 2024 · 被引用 453 次
- Provable Robust Watermarking for AI-Generated TextXuandong Zhao, Prabhanjan Vijendra Ananth, Lei Li, Yu-Xiang WangICLR 2024 · 被引用 312 次
- Compact Language Models via Pruning and Knowledge DistillationSaurav Muralidharan, Sharath Turuvekere Sreenivas, Raviraj Joshi, Marcin Chochowski 等NeurIPS 2024 · 被引用 198 次
相关 Paper
- Model Equality Testing: Which Model is this API Serving?Irena Gao, Percy Liang, Carlos GuestrinICLR 2025
- Sparse Autoencoders Trained on the Same Data Learn Different FeaturesGonçalo Paulo, Nora BelroseICLR 2026 · 被引用 96 次
- Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMsZiqian Zhong, Aditi RaghunathanICLR 2026 · 被引用 7 次
- AWM: Accurate Weight-Matrix Fingerprint for Large Language ModelsBoyi Zeng, Lin Chen, Ziwei He, Xinbing Wang 等ICLR 2026 · 被引用 3 次
- REEF: Representation Encoding Fingerprints for Large Language ModelsJie Zhang, Dongrui Liu, Chen Qian, Linfeng Zhang 等ICLR 2025
