Echoes within the Reasoning: Stealthy and Effective Watermarking via Chain of Thought
Jiacheng Lu, Yiming Li, Tao Song, Weijian Wang, Wenjie Qu, Haibing Guan, Jiaheng Zhang
摘要
Large Language Models with Chain-of-Thought reasoning capabilities represent valuable intellectual property, yet existing black-box watermarking methods often trade robustness for reasoning fidelity by perturbing final answers or relying on fragile trigger patterns. We propose BiCoT, a watermarking framework that embeds ownership signals into the internal geometry of reasoning traces by aligning high-saliency structural anchors with a private signature subspace while regularizing ordinary control tokens to preserve semantic capacity. This design couples the watermark with reasoning-relevant representations, making removal difficult without disrupting the features that support coherent reasoning. To enable verification under model theft and representation drift, we introduce Robust Subspace Registration (RSR), a Top-k logprob-based black-box verifier that uses sentinel tokens to calibrate systematic shifts in the output distribution. Experiments show that BiCoT preserves reasoning fidelity across diverse complex reasoning tasks while achieving robust detection under fine-tuning, quantization, model-level perturbations, and adaptive output-level attacks across in-domain and out-ofdistribution settings. Code is available at https: //github.com/JackLo111/BiCoT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper15
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- ABS: Scanning Neural Networks for Back-doors by Artificial Brain StimulationYingqi Liu, Wen-Chuan Lee, Guanhong Tao, Shiqing Ma 等CCS 2019 · 被引用 531 次
- Towards Reliable and Efficient Backdoor Trigger Inversion via Decoupling Benign FeaturesXiong Xu, Kunzhe Huang, Yiming Li, Zhan Qin 等ICLR 2024 · 被引用 59 次
相关 Paper
- Distilling the Thought, Watermarking the Answer: A Principle Semantic Guided Watermark for Reasoning Large Language ModelsShuliang Liu, Xingyu Li, Hongyi Liu, Dong Fang 等ICLR 2026 · 被引用 2 次
- ReasMark: A Robust Watermark for Attributing LLM Reasoning Under Knowledge Distillation AttacksPeizhuo Lv, Ruihua Zhou, Yunpeng Li, Ruigang Liang 等ACL 2026
- Verifying Chain-of-Thought Reasoning via Its Computational GraphZheng Zhao, Yeskendir Koishekenov, Xianjun Yang, Naila Murray 等ICLR 2026 · 被引用 25 次
- ImF: Embedding an Implicit Fingerprint in Your Large Language ModelsJiaxuan Wu, Wanli Peng, Hang Fu, Yiming Xue 等ACL 2026
- CTCC: A Robust and Stealthy Fingerprinting Framework for Large Language Models via Cross-Turn Contextual Correlation BackdoorZhenhua Xu, Xixiang Zhao, Xubin Yue, Shengwei Tian 等EMNLP 2025
