Zero-Shot Detection of LLM-Generated Text via Implicit Reward Model
Runheng Liu, Heyan Huang, Xingchen Xiao, Zhijing Wu
摘要
Large language models (LLMs) have demonstrated remarkable capabilities across various tasks. However, their ability to generate human-like text has raised concerns about potential misuse. This underscores the need for reliable and effective methods to detect LLM-generated text. In this paper, we propose IRM, a novel zero-shot approach that leverages Implicit Reward Models for LLM-generated text detection. Such implicit reward models can be derived from publicly available instruction-tuned and base models. Previous reward-based method relies on preference construction and task-specific fine-tuning. In comparison, IRM requires neither preference collection nor additional training. We evaluate IRM on the DetectRL benchmark and demonstrate that IRM can achieve superior detection performance, outperforms existing zero-shot and supervised methods in LLM-generated text detection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Watermarking Diffusion Language ModelsThibaud Gloaguen, Robin Staab, Nikola Jovanović, Martin VechevICLR 2026 · 被引用 13 次
- Zero-Shot Detection of LLM-Generated Text using Temperature SensitivityShixuan Ma, Jiahao Li, Zhendong Mao, Quan WangACL 2026
- Exons-Detect: Identifying and Amplifying Exonic Tokens via Hidden-State Discrepancy for Robust AI-Generated Text DetectionXiaowei Zhu, Yubing Ren, Fang Fang, Shi Wang 等ACL 2026
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning 等ICML 2023 · 被引用 988 次
- RADAR: Robust AI-Text Detection via Adversarial LearningXiaomeng Hu, Pin-Yu Chen, Tsung-Yi HoNeurIPS 2023 · 被引用 315 次
相关 Paper
- ReMoDetect: Reward Models Recognize Aligned LLM's GenerationsHyunseok Lee, Jihoon Tack, Jinwoo ShinNeurIPS 2024 · 被引用 13 次
- Unveiling the Implicit Toxicity in Large Language ModelsJiaxin Wen, Pei Ke, Hao Sun, Zhexin Zhang 等EMNLP 2023 · 被引用 21 次
- DNA-DetectLLM: Unveiling AI-Generated Text via a DNA-Inspired Mutation-Repair ParadigmXiaowei Zhu, Yubing Ren, Fang Fang, Qingfeng Tan 等NeurIPS 2025 · 被引用 10 次
- HLD: Approximate Hierarchical Linguistic Distribution Modeling for LLM-Generated Text DetectionRui Guo, Weibin Zeng, Fuzhang Wu, Yan Kong 等ICLR 2026
- Uncovering LLM-Generated Code: A Zero-Shot Synthetic Code Detector via Code RewritingTong Ye, Yangkai Du, Tengfei Ma, Lingfei Wu 等AAAI 2025 · 被引用 21 次
