Ideal Attribution and Faithful Watermarks for Language Models
Min Jae Song, Kameron Shahabi
Abstract
We introduce ideal attribution mechanisms, a formal abstraction for reasoning about attribution decisions over strings. At the core of this abstraction lies the ledger, an append-only log of the prompt-response interaction history between a model and its user. Each mechanism produces deterministic decisions based on the ledger and an explicit selection criterion, making it well-suited to serve as a ground truth for attribution. We frame the design goal of watermarking schemes as faithful representation of ideal attribution mechanisms. This novel perspective brings conceptual clarity, replacing piecemeal probabilistic statements with a unified language for stating the guarantees of each scheme. It also enables precise reasoning about desiderata for future watermarking schemes, even when no current construction achieves them, since the ideal functionalities are specified first. In this way, the framework provides a roadmap that clarifies which guarantees are attainable in an idealized setting and worth pursuing in practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 45cbd9db-ee59-4752-bbd3-7ab803a0c782Builds on9
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
- Provable Robust Watermarking for AI-Generated TextXuandong Zhao, Prabhanjan Vijendra Ananth, Lei Li, Yu-Xiang WangICLR 2024 · 312 citations
- On the Reliability of Watermarks for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu et al.ICLR 2024 · 202 citations
- Pseudorandom Error-Correcting CodesMiranda Christ, Sam GunnCRYPTO 2024 · 17 citations
- New Perspectives on the Polyak Stepsize: Surrogate Functions and Negative ResultsFrancesco Orabona, Ryan D'OrazioNeurIPS 2025 · 9 citations
Related papers
- ProMark: Proactive Diffusion Watermarking for Causal AttributionVishal Asnani, John P. Collomosse, Tu Bui, Xiaoming Liu et al.CVPR 2024
- Anchor Watermark: Robust Attribution for Diffusion-based Text-to-Audio ModelXianjin Rong, Donghui HuAAAI 2026
- Watermark-based Attribution of AI-Generated ContentZhengyuan Jiang, Moyang Guo, Yuepeng Hu, Yupu Wang et al.ICLR 2026 · 11 citations
- Unforgeable Watermarks for Language Models via Robust SignaturesHuijia Lin, Kameron Shahabi, Min Jae SongCRYPTO 2026
- IACW: Intent-Aware Controllable Watermarking for Scalable Authorial Intent AttributionHao Huang, Ruihua Zhou, JiaTang Luo, Yunpeng Li et al.ICML 2026
