Lune

ICML2025顶会

Independence Tests for Language Models

Sally Zhu, Ahmed M. Ahmed, Rohith Kuditipudi, Percy Liang

出版方
2025年份
7顶会引用

摘要

We consider the following problem: given the weights of two models, can we test whether they were trained independently-i.e., from independent random initializations? We consider two settings: constrained and unconstrained. In the constrained setting, we make assumptions about model architecture and training and propose a family of statistical tests that yield exact p-values with respect to the null hypothesis that the models are trained from independent random initializations. These p-values are valid regardless of the composition of either model's training data; we compute them by simulating exchangeable copies of each model under our assumptions and comparing various similarity measures of weights and activations between the original two models versus these copies. We report the p-values from these tests on pairs of 21 open-weight models (210 total pairs) and find we correctly identify all pairs of non-independent models. Notably, our tests remain effective even if one of the models was fine-tuned for many tokens. In the unconstrained setting, where we make no assumptions about training procedures, can change model architecture, and allow for adversarial evasion attacks, the previous tests no longer work. Instead, we propose a new test which matches hidden activations between two models, and use it to construct a test that is robust to adversarial transformations and to changes in model architecture. The test can also perform localized testing: identifying specific non-independent components of models. Though we no longer obtain exact p-values from this test, empirically we find it behaves as one and reliably distinguishes nonindependent models. Notably, we can use the test to identify specific parts of one model that are derived from another (e.g., how Llama 3.1-8B was pruned to initialize Llama 3.2-3B, or shared layers between Mistral-7B and StripedHyena-7B), and it is even robust to retraining individual layers of either model from scratch.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 6752f2cc-d413-46b3-a3e4-bccc6ed762b2

引用它的顶会 Paper7

问问它们各自怎么用它

它引用的顶会 Paper10

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖