LORD-GoF: A Robust Online Detection Approach for LLM Watermarks in Sparse and Mixed Streams
Jiade Xu, Zhouping Li
Abstract
Watermarking is crucial for identifying AIgenerated text; however, existing detection methods often focus on offline settings and fail to control the online False Discovery Rate (oFDR) when applied to real-world streams where machinegenerated content is sparse and mixed with human writing. To address this issue, in this paper, we propose LORD-GOF, a novel online detection framework that combines a Goodness-of-Fit (GoF) statistic with the Levels based On Recent Discovery (LORD) procedure. We prove that the LORD-GOF approach can rigorously control the oFDR below a user-specified level by dynamically adjusting detection thresholds. Extensive experiments on watermarked text from Qwen-2.5-3B, Sheared-LLaMA-2.7B, and OPT-1.3B using both the Gumbel-Max and Inverse Transform watermarking schemes show that our method maintains statistical power comparable to offline benchmarks while successfully controlling the oFDR under complex, mixed streaming scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4894cb5e-47dd-4290-867d-ff0ada2ede0dBuilds on7
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
- Watermark Stealing in Large Language ModelsNikola Jovanovic, Robin Staab, Martin T. VechevICML 2024 · 88 citations
- Adaptive Text Watermark for Large Language ModelsYepeng Liu, Yuheng BuICML 2024 · 63 citations
- Token-Specific Watermarking with Enhanced Detectability and Semantic Coherence for Large Language ModelsMingjia Huo, Sai Ashish Somayajula, Youwei Liang, Ruisi Zhang et al.ICML 2024 · 37 citations
- On the Empirical Power of Goodness-of-Fit Tests in Watermark DetectionWeiqing He, Xiang Li, Tianqi Shang, Li Shen et al.NeurIPS 2025 · 6 citations
Related papers
- Efficiently Identifying Watermarked Segments in Mixed-Source TextsXuandong Zhao, Chenwen Liao, Yuxiang Wang, Lei LiACL 2025 · 3 citations
- On the Reliability of Watermarks for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Manli Shu et al.ICLR 2024 · 202 citations
- Online Detection of LLM-Generated Texts via Sequential Hypothesis Testing by BettingCan Chen, Jun-Kun WangICML 2025
- How Good is Post-Hoc Watermarking With Language Model Rephrasing?Pierre Fernandez, Tom Sander, Hady Elsahar, Hongyan Chang et al.ICML 2026 · 2 citations
- GumbelSoft: Diversified Language Model Watermarking via the GumbelMax-trickJiayi Fu, Xuandong Zhao, Ruihan Yang, Yuansen Zhang et al.ACL 2024 · 6 citations
