DetectAnyLLM: Towards Generalizable and Robust Detection of Machine-Generated Text Across Domains and Models
Jiachen Fu, Chun-Le Guo, Chongyi Li
Abstract
The rapid advancement of large language models (LLMs) has drawn urgent attention to the task of machine-generated text detection (MGTD). However, existing approaches struggle in complex real-world scenarios: zero-shot detectors rely heavily on scoring model's output distribution while training-based detectors are often constrained by overfitting to the training data, limiting generalization. We found that the performance bottleneck of training-based detectors stems from the misalignment between training objective and task needs. To address this, we propose Direct Discrepancy Learning (DDL), a novel optimization strategy that directly optimizes the detector with task-oriented knowledge. DDL enables the detector to better capture the core semantics of the detection task, thereby enhancing both robustness and generalization. Built upon this, we introduce DetectAnyLLM, a unified detection framework that achieves state-of-the-art MGTD performance across diverse LLMs. To ensure a reliable evaluation, we construct MIRAGE, the most diverse multi-task MGTD benchmark. MIRAGE samples human-written texts from 10 corpora across 5 text-domains, which are then re-generated or revised using 17 cutting-edge LLMs, covering a wide spectrum of proprietary models and textual styles. Extensive experiments on MIRAGE reveal the limitations of existing methods in complex environment. In contrast, DetectAnyLLM consistently outperforms them, achieving over a 70% performance improvement under the same training data and base scoring model, underscoring the effectiveness of our DDL. Project page: https://fjc2005.github.io/detectanyllm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dd815b93-d67b-4a18-b507-4ef948d9d16dCited by top-tier papers4
- Minimizing Mismatch Risk: A Prototype-Based Routing Framework for Zero-shot LLM-generated Text DetectionKe Sun, Guangsheng Bao, Han Cui, Yue ZhangICML 2026 · 1 citation
- Verifiable LLM-Generated Text Detection via Projected Semantic-Structural DistributionsRuochong Xiong, Qien Li, Wangwang Lian, Yulong Wan et al.ACL 2026
- HLD: Approximate Hierarchical Linguistic Distribution Modeling for LLM-Generated Text DetectionRui Guo, Weibin Zeng, Fuzhang Wu, Yan Kong et al.ICLR 2026
- Robust Membership Inference for Large Language Models under Adversarial Generative CorruptionYuanhong Huang, Huili Wang, Xueying Bai, Jinrui Wang et al.ACL 2026
Builds on19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning et al.ICML 2023 · 988 citations
- RADAR: Robust AI-Text Detection via Adversarial LearningXiaomeng Hu, Pin-Yu Chen, Tsung-Yi HoNeurIPS 2023 · 315 citations
- Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability CurvatureGuangsheng Bao, Yanbin Zhao, Zhiyang Teng, Linyi Yang et al.ICLR 2024 · 311 citations
Related papers
- Breaking the Generator Barrier: Disentangled Representation for Generalizable AI-Text DetectionXiao Pu, Zepeng Cheng, Lin Yuan, Yu Wu et al.ACL 2026 · 1 citation
- MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection BenchmarkDominik Macko, Róbert Móro, Adaku Uchendu, Jason Samuel Lucas et al.EMNLP 2023 · 25 citations
- DeTeCtive: Detecting AI-generated Text via Multi-Level Contrastive LearningXun Guo, Yongxin He, Shan Zhang, Ting Zhang et al.NeurIPS 2024 · 100 citations
- TSM-Bench: Detecting LLM-Generated Text in Real-World Wikipedia Editing PracticesGerrit Quaremba, Elizabeth Black, Denny Vrandecic, Elena SimperlICLR 2026 · 2 citations
- Detecting Machine-Generated Texts by Multi-Population Aware Optimization for Maximum Mean DiscrepancyShuhai Zhang, Yiliao Song, Jiahao Yang, Yuanqing Li et al.ICLR 2024 · 19 citations
