Matching Pairs: Attributing Fine-Tuned Models to their Pre-Trained Large Language Models
Myles Foley, Ambrish Rawat, Taesung Lee, Yufang Hou, Gabriele Picco, Giulio Zizzo
Abstract
The wide applicability and adaptability of generative large language models (LLMs) has enabled their rapid adoption. While the pretrained models can perform many tasks, such models are often fine-tuned to improve their performance on various downstream applications. However, this leads to issues over violation of model licenses, model theft, and copyright infringement. Moreover, recent advances show that generative technology is capable of producing harmful content which exacerbates the problems of accountability within model supply chains. Thus, we need a method to investigate how a model was trained or a piece of text was generated and what their pre-trained base model was. In this paper we take the first step to address this open problem by tracing back the origin of a given fine-tuned LLM to its corresponding pre-trained base model. We consider different knowledge levels and attribution strategies, and find that we can correctly trace back 8 out of the 10 fine tuned models with our best method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- RealCompo: Balancing Realism and Compositionality Improves Text-to-Image Diffusion ModelsXinchen Zhang, Ling Yang, Yaqi Cai, Zhaochen Yu et al.NeurIPS 2024 · 22 citations
- Safety Alignment via Constrained Knowledge UnlearningZesheng Shi, Yucheng Zhou, Jing Li, Yuxin Jin et al.ACL 2025 · 8 citations
Builds on8
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisErik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu et al.ICLR 2023 · 234 citations
- Recall and Learn: Fine-tuning Deep Pretrained Language Models with Less ForgettingSanyuan Chen, Yutai Hou, Yiming Cui, Wanxiang Che et al.EMNLP 2020 · 152 citations
Related papers
- Continual Origin Tracing of LLM-Generated TextHaoran Li, Quan WangSIGIR 2025 · 2 citations
- Tracking the Copyright of Large Vision-Language Models through Parameter Learning Adversarial ImagesYubo Wang, Jianting Tang, Chaohu Liu, Linli XuICLR 2025
- Model Provenance Testing for Large Language ModelsIvica Nikolic, Teodora Baluta, Prateek SaxenaNeurIPS 2025 · 20 citations
- Identifying Provenance of Generative Text-to-Image ModelsAnna Yoo Jeong Ha, Wenxin Ding, Stanley Wu, Shawn Shan et al.USENIX Security 2026
- HuRef: HUman-REadable Fingerprint for Large Language ModelsBoyi Zeng, Lizheng Wang, Yuncong Hu, Yi Xu et al.NeurIPS 2024 · 48 citations
