Independence Tests for Language Models
Sally Zhu, Ahmed M. Ahmed, Rohith Kuditipudi, Percy Liang
Abstract
We consider the following problem: given the weights of two models, can we test whether they were trained independently-i.e., from independent random initializations? We consider two settings: constrained and unconstrained. In the constrained setting, we make assumptions about model architecture and training and propose a family of statistical tests that yield exact p-values with respect to the null hypothesis that the models are trained from independent random initializations. These p-values are valid regardless of the composition of either model's training data; we compute them by simulating exchangeable copies of each model under our assumptions and comparing various similarity measures of weights and activations between the original two models versus these copies. We report the p-values from these tests on pairs of 21 open-weight models (210 total pairs) and find we correctly identify all pairs of non-independent models. Notably, our tests remain effective even if one of the models was fine-tuned for many tokens. In the unconstrained setting, where we make no assumptions about training procedures, can change model architecture, and allow for adversarial evasion attacks, the previous tests no longer work. Instead, we propose a new test which matches hidden activations between two models, and use it to construct a test that is robust to adversarial transformations and to changes in model architecture. The test can also perform localized testing: identifying specific non-independent components of models. Though we no longer obtain exact p-values from this test, empirically we find it behaves as one and reliably distinguishes nonindependent models. Notably, we can use the test to identify specific parts of one model that are derived from another (e.g., how Llama 3.1-8B was pruned to initialize Llama 3.2-3B, or shared layers between Mistral-7B and StripedHyena-7B), and it is even robust to retraining individual layers of either model from scratch.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6752f2cc-d413-46b3-a3e4-bccc6ed762b2Cited by top-tier papers7
- LLM DNA: Tracing Model Evolution via Functional RepresentationsZhaomin Wu, Haodong Zhao, Ziyang Wang, Jizhou Guo et al.ICLR 2026 · 21 citations
- Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model OutputsYiwei Chen, Soumyadeep Pal, Yimeng Zhang, Qing Qu et al.ICLR 2026 · 15 citations
- Blackbox Model Provenance via Palimpsestic Membership InferenceRohith Kuditipudi, Jing Huang, Sally Zhu, Diyi Yang et al.NeurIPS 2025 · 12 citations
- ZipLLM: Efficient LLM Storage via Model-Aware Synergistic Data Deduplication and CompressionZirui Wang, Tingfeng Lan, Zhaoyuan Su, Juncheng Yang et al.NSDI 2026 · 8 citations
- Adaptive Multiscale Binary Expansion Tests for IndependenceYang Yang, Duo Zheng, Sandeep Jain, Kai Zhang et al.ICML 2026
Builds on10
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin et al.NeurIPS 2023 · 1,975 citations
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
- Sheared LLaMA: Accelerating Language Model Pre-training via Structured PruningMengzhou Xia, Tianyu Gao, Zhiyuan Zeng, Danqi ChenICLR 2024 · 453 citations
- Provable Robust Watermarking for AI-Generated TextXuandong Zhao, Prabhanjan Vijendra Ananth, Lei Li, Yu-Xiang WangICLR 2024 · 312 citations
- Compact Language Models via Pruning and Knowledge DistillationSaurav Muralidharan, Sharath Turuvekere Sreenivas, Raviraj Joshi, Marcin Chochowski et al.NeurIPS 2024 · 198 citations
Related papers
- Model Equality Testing: Which Model is this API Serving?Irena Gao, Percy Liang, Carlos GuestrinICLR 2025
- Sparse Autoencoders Trained on the Same Data Learn Different FeaturesGonçalo Paulo, Nora BelroseICLR 2026 · 96 citations
- Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMsZiqian Zhong, Aditi RaghunathanICLR 2026 · 7 citations
- AWM: Accurate Weight-Matrix Fingerprint for Large Language ModelsBoyi Zeng, Lin Chen, Ziwei He, Xinbing Wang et al.ICLR 2026 · 3 citations
- REEF: Representation Encoding Fingerprints for Large Language ModelsJie Zhang, Dongrui Liu, Chen Qian, Linfeng Zhang et al.ICLR 2025
