Barriers to Counterfactual Credit Attribution for Autoregressive Models
Aloni Cohen, Chenhao Zhang
Abstract
Generative AI disrupts the practice of giving credit to work that came before. Ideally, a generative model would give credit to any work on which its output depends in a significant way. Counterfactual credit attribution (CCA) is a technical condition formalizing this goal-a relaxation of differential privacy-recently introduced by Livni, Moran, Nissim, and Pabbaraju [2024] who studied it in the PAC learning setting. We initiate the study of CCA generative models. Specifically, we consider autoregressive models giving credit to a deployment-time dataset (e.g., a RAG database). We uncover barriers to two natural approaches to CCA autoregressive models. First, we show that imposing CCA on the underlying next-token predictor does not guarantee that the model is CCA: CCA does not compose autoregressively (unlike DP). Second, we consider a different approach to building CCA models which we call retrofitting. Retrofitting takes a model that does not attribute credit, and adds credit onto it. We prove a lower bound for CCA retrofitting under a weak optimality requirement. Given black-box access to the starting model, retrofitting requires query complexity exponential in the length of the model's outputs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 288da5b0-f2e6-4331-bce2-b7df6003236eBuilds on7
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- On Provable Copyright Protection for Generative ModelsNikhil Vyas, Sham M. Kakade, Boaz BarakICML 2023 · 120 citations
- Privacy-Preserving In-Context Learning with Differentially Private Few-Shot GenerationXinyu Tang, Richard Shin, Huseyin A. Inan, Andre Manoel et al.ICLR 2024 · 111 citations
- Private Reinforcement Learning with PAC and Regret GuaranteesGiuseppe Vietri, Borja Balle, Akshay Krishnamurthy, Zhiwei Steven WuICML 2020 · 70 citations
- Black-Box Differential Privacy for Interactive MLHaim Kaplan, Yishay Mansour, Shay Moran, Kobbi Nissim et al.NeurIPS 2023 · 7 citations
Related papers
- Credit Attribution and Stable CompressionRoi Livni, Shay Moran, Kobbi Nissim, Chirag PabbarajuNeurIPS 2024 · 4 citations
- IncreFA: Breaking the Static Wall of Generative Model AttributionHaotian Qin, Dongliang Chang, Yueying Gao, Yuexuan Tan et al.CVPR 2026 · 3 citations
- CPR: Retrieval Augmented Generation for Copyright ProtectionAditya Golatkar, Alessandro Achille, Luca Zancato, Yu-Xiang Wang et al.CVPR 2024
- Blameless Users in a Clean Room: Defining Copyright Protection for Generative ModelsAloni CohenNeurIPS 2025 · 2 citations
- Auditing Private PredictionKaran Chadha, Matthew Jagielski, Nicolas Papernot, Christopher A. Choquette-Choo et al.ICML 2024 · 10 citations
