Zero-Shot Detection of LLM-Generated Text using Token Cohesiveness
Shixuan Ma, Quan Wang
摘要
The increasing capability and widespread usage of large language models (LLMs) highlight the desirability of automatic detection of LLMgenerated text. Zero-shot detectors, due to their training-free nature, have received considerable attention and notable success. In this paper, we identify a new feature, token cohesiveness, that is useful for zero-shot detection, and we demonstrate that LLM-generated text tends to exhibit higher token cohesiveness than human-written text. Based on this observation, we devise TOC-SIN, a generic dual-channel detection paradigm that uses token cohesiveness as a plug-and-play module to improve existing zero-shot detectors. To calculate token cohesiveness, TOCSIN only requires a few rounds of random token deletion and semantic difference measurement, making it particularly suitable for a practical black-box setting where the source model used for generation is not accessible. Extensive experiments with four state-of-the-art base detectors on various datasets, source models, and evaluation settings demonstrate the effectiveness and generality of the proposed approach. Code available at: https://github.com/Shixuan-Ma/TOCSIN .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Beyond Binary: Towards Fine-Grained LLM-Generated Text Detection via Role Recognition and Involvement MeasurementZihao Cheng, Li Zhou, Feng Jiang, Benyou Wang 等WWW 2025 · 被引用 20 次
- Watermarking Diffusion Language ModelsThibaud Gloaguen, Robin Staab, Nikola Jovanović, Martin VechevICLR 2026 · 被引用 13 次
- Zero-Shot Detection of LLM-Generated Text via Implicit Reward ModelRunheng Liu, Heyan Huang, Xingchen Xiao, Zhijing WuNeurIPS 2025 · 被引用 7 次
- Beyond Raw Detection Scores: Markov-Informed Calibration for Boosting Machine-Generated Text DetectionChenwang Wu, Yiu-ming Cheung, Shuhai Zhang, Bo Han 等ICLR 2026 · 被引用 2 次
- Advancing Machine-Generated Text Detection from an Easy to Hard Supervision PerspectiveChenwang Wu, Yiu-ming Cheung, Bo Han, Defu LianNeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper8
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 被引用 1,143 次
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning 等ICML 2023 · 被引用 988 次
- Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability CurvatureGuangsheng Bao, Yanbin Zhao, Zhiyang Teng, Linyi Yang 等ICLR 2024 · 被引用 311 次
- DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated TextXianjun Yang, Wei Cheng, Yue Wu, Linda Ruth Petzold 等ICLR 2024 · 被引用 173 次
相关 Paper
- Zero-Shot Detection of LLM-Generated Text using Temperature SensitivityShixuan Ma, Jiahao Li, Zhendong Mao, Quan WangACL 2026
- Telescope: Improving Zero Shot Detection of LLM Generated Content By Measuring Token Repetition ProbabilityChristopher Nassif, Joshua CooperICML 2026
- Training-free LLM-generated Text Detection by Mining Token Probability SequencesYihuai Xu, Yongwei Wang, Yifei Bi, Huangsen Cao 等ICLR 2025
- HLD: Approximate Hierarchical Linguistic Distribution Modeling for LLM-Generated Text DetectionRui Guo, Weibin Zeng, Fuzhang Wu, Yan Kong 等ICLR 2026
- DualCodeDetect: Zero-Shot LLM-Generated Code Detection via Dual-Channel PerturbationZhengdao Li, Xiuwei Shang, Zhenkan Fu, Shikai Guo 等FSE 2026
