SeedPrints: Fingerprints Can Even Tell Which Seed Your Large Language Model Was Trained From
Yao Tong, Haonan Wang, Siquan Li, Kenji Kawaguchi, Tianyang Hu
Abstract
Fingerprinting Large Language Models (LLMs) is essential for provenance verification and model attribution. Existing fingerprinting methods are primarily evaluated after fine-tuning, where models have already acquired stable signatures from training data, optimization dynamics, or hyperparameters. However, most of a model's capacity and knowledge are acquired during pretraining rather than downstream fine-tuning, making large-scale pretraining a more fundamental regime for lineage verification. We show that existing fingerprinting methods become unreliable in this regime, as they rely on post-hoc signatures that only emerge after substantial training. This limitation contradicts the classical Galton notion of a fingerprint as an intrinsic and persistent identity. In contrast, we propose a stronger and more intrinsic notion of LLM fingerprinting: SeedPrints, a method that leverages random initialization biases as persistent, seed-dependent identifiers present even before training begins. We show that untrained models exhibit reproducible prediction biases induced by their initialization seed, and that these weak signals remain statistically detectable throughout training, enabling high-confidence lineage verification. Unlike prior techniques that fail during early pretraining or degrade under distribution shifts, SeedPrints remains effective across all training stages, from initialization to large-scale pretraining and downstream adaptation. Experiments on LLaMA-style and Qwen-style models demonstrate seed-level distinguishability and enable birth-to-lifecycle identity verification. Evaluations on large-scale pretraining trajectories and real-world fingerprinting benchmarks further confirm its robustness under prolonged training, domain shifts, and parameter modifications. Together, our results show that initialization itself imprints a unique and persistent identity on LLMs, forming a true "Galtonian" fingerprint. Code is available at https://github.com/YnezT0311/SeedPrints .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6b466166-0e98-470c-bbb4-abf63410896eCited by top-tier papers1
Ask how each one uses itBuilds on9
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song et al.S&P 2022 · 1,049 citations
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
- Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by BackdooringYossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas et al.USENIX Security 2018 · 832 citations
- Llemma: An Open Language Model for MathematicsZhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos et al.ICLR 2024 · 433 citations
- Watermarking Makes Language Models RadioactiveTom Sander, Pierre Fernandez, Alain Durmus, Matthijs Douze et al.NeurIPS 2024 · 68 citations
Related papers
- Fingerprinting LLMs via Prompt InjectionYuepeng Hu, Zhengyuan Jiang, Mengyuan Li, Osama Ahmed et al.ACL 2026 · 3 citations
- Ghost in the Transformer: Detecting Model Reuse with Invariant Spectral SignaturesSuqing Wang, Ziyang Ma, Xinyi Li, Zuchao LiAAAI 2026 · 1 citation
- AWM: Accurate Weight-Matrix Fingerprint for Large Language ModelsBoyi Zeng, Lin Chen, Ziwei He, Xinbing Wang et al.ICLR 2026 · 3 citations
- HuRef: HUman-REadable Fingerprint for Large Language ModelsBoyi Zeng, Lizheng Wang, Yuncong Hu, Yi Xu et al.NeurIPS 2024 · 48 citations
- Matching Pairs: Attributing Fine-Tuned Models to their Pre-Trained Large Language ModelsMyles Foley, Ambrish Rawat, Taesung Lee, Yufang Hou et al.ACL 2023 · 2 citations
