Barriers to Counterfactual Credit Attribution for Autoregressive Models
Aloni Cohen, Chenhao Zhang
摘要
Generative AI disrupts the practice of giving credit to work that came before. Ideally, a generative model would give credit to any work on which its output depends in a significant way. Counterfactual credit attribution (CCA) is a technical condition formalizing this goal-a relaxation of differential privacy-recently introduced by Livni, Moran, Nissim, and Pabbaraju [2024] who studied it in the PAC learning setting. We initiate the study of CCA generative models. Specifically, we consider autoregressive models giving credit to a deployment-time dataset (e.g., a RAG database). We uncover barriers to two natural approaches to CCA autoregressive models. First, we show that imposing CCA on the underlying next-token predictor does not guarantee that the model is CCA: CCA does not compose autoregressively (unlike DP). Second, we consider a different approach to building CCA models which we call retrofitting. Retrofitting takes a model that does not attribute credit, and adds credit onto it. We prove a lower bound for CCA retrofitting under a weak optimality requirement. Given black-box access to the starting model, retrofitting requires query complexity exponential in the length of the model's outputs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- On Provable Copyright Protection for Generative ModelsNikhil Vyas, Sham M. Kakade, Boaz BarakICML 2023 · 被引用 120 次
- Privacy-Preserving In-Context Learning with Differentially Private Few-Shot GenerationXinyu Tang, Richard Shin, Huseyin A. Inan, Andre Manoel 等ICLR 2024 · 被引用 111 次
- Private Reinforcement Learning with PAC and Regret GuaranteesGiuseppe Vietri, Borja Balle, Akshay Krishnamurthy, Zhiwei Steven WuICML 2020 · 被引用 70 次
- Black-Box Differential Privacy for Interactive MLHaim Kaplan, Yishay Mansour, Shay Moran, Kobbi Nissim 等NeurIPS 2023 · 被引用 7 次
相关 Paper
- Credit Attribution and Stable CompressionRoi Livni, Shay Moran, Kobbi Nissim, Chirag PabbarajuNeurIPS 2024 · 被引用 4 次
- IncreFA: Breaking the Static Wall of Generative Model AttributionHaotian Qin, Dongliang Chang, Yueying Gao, Yuexuan Tan 等CVPR 2026 · 被引用 3 次
- CPR: Retrieval Augmented Generation for Copyright ProtectionAditya Golatkar, Alessandro Achille, Luca Zancato, Yu-Xiang Wang 等CVPR 2024
- Blameless Users in a Clean Room: Defining Copyright Protection for Generative ModelsAloni CohenNeurIPS 2025 · 被引用 2 次
- Auditing Private PredictionKaran Chadha, Matthew Jagielski, Nicolas Papernot, Christopher A. Choquette-Choo 等ICML 2024 · 被引用 10 次
