ReMoDetect: Reward Models Recognize Aligned LLM's Generations
Hyunseok Lee, Jihoon Tack, Jinwoo Shin
Abstract
The remarkable capabilities and easy accessibility of large language models (LLMs) have significantly increased societal risks (e.g., fake news generation), necessitating the development of LLM-generated text (LGT) detection methods for safe usage. However, detecting LGTs is challenging due to the vast number of LLMs, making it impractical to account for each LLM individually; hence, it is crucial to identify the common characteristics shared by these models. In this paper, we draw attention to a common feature of recent powerful LLMs, namely the alignment training, i.e., training LLMs to generate human-preferable texts. Our key finding is that as these aligned LLMs are trained to maximize the human preferences, they generate texts with higher estimated preferences even than human-written texts; thus, such texts are easily detected by using the reward model (i.e., an LLM trained to model human preference distribution). Based on this finding, we propose two training schemes to further improve the detection ability of the reward model, namely (i) continual preference fine-tuning to make the reward model prefer aligned LGTs even further and (ii) reward modeling of Human/LLM mixed texts (a rephrased texts from human-written texts using aligned LLMs), which serves as a median preference text corpus between LGTs and human-written texts to learn the decision boundary better. We provide an extensive evaluation by considering six text domains across twelve aligned LLMs, where our method demonstrates state-of-the-art results. Code is available at https://github.com/hyunseoklee-ai/ReMoDetect.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d4853322-4bbb-4bb0-b147-906ff8a00cabCited by top-tier papers7
- Learn-to-Distance: Distance Learning for Detecting LLM-Generated TextHongyi Zhou, Jin Zhu, Kai Ye, Ying Yang et al.ICLR 2026 · 10 citations
- Zero-Shot Detection of LLM-Generated Text via Implicit Reward ModelRunheng Liu, Heyan Huang, Xingchen Xiao, Zhijing WuNeurIPS 2025 · 7 citations
- Advancing Machine-Generated Text Detection from an Easy to Hard Supervision PerspectiveChenwang Wu, Yiu-ming Cheung, Bo Han, Defu LianNeurIPS 2025 · 2 citations
- Attacks on Machine-Text Detectors Retain Stylistic FingerprintsRafael Rivera Soto, Barry Chen, Nicholas AndrewsICML 2026 · 1 citation
- LLMs Killed Q&A Stars? Analyzing the Impact of LLM-Generated Answers on an Online Q&A PlatformDongwon Shin, Sooel SonWWW 2026
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 2,230 citations
- AugMix: A Simple Data Processing Method to Improve Robustness and UncertaintyDan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph et al.ICLR 2020 · 1,572 citations
Related papers
- HLPD: Aligning LLMs to Human Language Preference for Machine-Revised Text DetectionFangqi Dai, Xingjian Jiang, Zizhuang DengAAAI 2026 · 1 citation
- MAGE: Machine-generated Text Detection in the WildYafu Li, Qintong Li, Leyang Cui, Wei Bi et al.ACL 2024 · 44 citations
- Pretraining Language Models with Human PreferencesTomasz Korbak, Kejian Shi, Angelica Chen, Rasika Vinayak Bhalerao et al.ICML 2023 · 287 citations
- Mission Impossible: A Statistical Perspective on Jailbreaking LLMsJingtong Su, Julia Kempe, Karen UllrichNeurIPS 2024 · 38 citations
- Imitate Before Detect: Aligning Machine Stylistic Preference for Machine-Revised Text DetectionJiaqi Chen, Xiaoye Zhu, Tianyang Liu, Ying Chen et al.AAAI 2025 · 13 citations
