AliMark: Enhancing Robustness of Sentence-Level Watermarking Against Text Paraphrasing
Yuexin Li, Wenjie Qu, Linyu Wu, Yulin Chen, Yufei He, Tri Cao, Bryan Hooi, Jiaheng Zhang
摘要
Existing sentence-level watermarking methods enhance robustness to paraphrasing by anchoring watermarks in sentence semantics. However, their prefix-based designs remain vulnerable to structural perturbations, such as sentence splitting and merging, which commonly arise under strong paraphrasers like DIPPER and GPT-3.5. To mitigate this issue, we propose AliMark, a framework that reformulates sentence-level watermarking as a bit sequence encoding and alignment problem between a potentially watermarked text and a secret bit sequence. Notably, our approach adopts a two-stage detection strategy: we generate multiple restructured text variants and adaptively align their extracted bit sequences with the secret bit sequence to minimize alignment cost. This multi-candidate alignment design naturally improves robustness to sentence merges and splits. Extensive experiments demonstrate that Al-iMark substantially outperforms state-of-the-art baselines under diverse paraphrasing attacks. Our code is available at https://github.com/ imethanlee/AliMark .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning 等ICML 2023 · 被引用 988 次
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz 等ICML 2023 · 被引用 854 次
相关 Paper
- SimMark: A Robust Sentence-Level Similarity-Based Watermarking Algorithm for Large Language ModelsAmirHossein Dabiri Aghdam, Lele WangEMNLP 2025 · 被引用 3 次
- Robust Multi-bit Text Watermark with LLM-based ParaphrasersXiaojun Xu, Jinghan Jia, Yuanshun Yao, Yang Liu 等ICML 2025
- PMark: Towards Robust and Distortion-free Semantic-level Watermarking with Channel ConstraintsJiahao Huo, Shuliang Liu, Bin Wang, Junyan Zhang 等ICLR 2026 · 被引用 18 次
- Towards Reliable Marking and Verification of AI-Generated Text via Geometry-aware Sentence-level WatermarkingYubing Ren, Ping Guo, Yanan CaoICML 2026
- Context-aware Watermark with Semantic Balanced Green-red Lists for Large Language ModelsYuxuan Guo, Zhiliang Tian, Yiping Song, Tianlun Liu 等EMNLP 2024 · 被引用 3 次
