M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection
Yuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su, Artem Shelmanov, Akim Tsvigun, Osama Mohammed Afzal, Tarek Mahmoud, Giovanni Puccetti, Thomas Arnold, Alham Fikri Aji, Nizar Habash
摘要
The advent of Large Language Models (LLMs) has brought an unprecedented surge in machinegenerated text (MGT) across diverse channels. This raises legitimate concerns about its potential misuse and societal implications. The need to identify and differentiate such content from genuine human-generated text is critical in combating disinformation, preserving the integrity of education and scientific fields, and maintaining trust in communication. In this work, we address this problem by introducing a new benchmark based on a multilingual, multidomain, and multi-generator corpus of MGTs -M4GT-Bench. The benchmark is compiled of three tasks: (1) mono-lingual and multi-lingual binary MGT detection; (2) multi-way detection where one need to identify, which particular model generated the text; and (3) mixed humanmachine text detection, where a word boundary delimiting MGT from human-written content should be determined. On the developed benchmark, we have tested several MGT detection baselines and also conducted an evaluation of human performance. We see that obtaining good performance in MGT detection usually requires an access to the training data from the same domain and generators. The benchmark is available at https://github. com/mbzuai-nlp/M4GT-Bench. Source Human Parallel Data Total New test Domain Total= Upsample+ Parallel davinci-003 ChatGPT Cohere Dolly-v2 BLOOMz Machine GPT-4 OUTFOX 16
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- DETree: DEtecting Human-AI Collaborative Texts via Tree-Structured Hierarchical Representation LearningYongxin He, Shan Zhang, Yixuan Cao, Lei Ma 等NeurIPS 2025 · 被引用 13 次
- Authorship Attribution in Multilingual Machine-Generated TextsLucio La Cava, Dominik Macko, Róbert Móro, Ivan Srba 等ACL 2026 · 被引用 7 次
- Is Human-Like Text Liked by Humans? Multilingual Human Detection and Preference Against AIYuxia Wang, Rui Xing, Jonibek Mansurov, Giovanni Puccetti 等ACL 2026 · 被引用 3 次
- TSM-Bench: Detecting LLM-Generated Text in Real-World Wikipedia Editing PracticesGerrit Quaremba, Elizabeth Black, Denny Vrandecic, Elena SimperlICLR 2026 · 被引用 2 次
- On the Salience of Low-Probability Tokens for AI-Generated Text Detection: A Multiscale Uncertainty PerspectiveYikai Guo, Bin Wang, Xilai Fan, Wenjun Ke 等ICML 2026
它引用的顶会 Paper17
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning 等ICML 2023 · 被引用 988 次
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz 等ICML 2023 · 被引用 854 次
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting 等NeurIPS 2023 · 被引用 657 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
相关 Paper
- MGTBench: Benchmarking Machine-Generated Text DetectionXinlei He, Xinyue Shen, Zeyuan Chen, Michael Backes 等CCS 2024 · 被引用 30 次
- MAGE: Machine-generated Text Detection in the WildYafu Li, Qintong Li, Leyang Cui, Wei Bi 等ACL 2024 · 被引用 44 次
- Beyond Binary: Towards Fine-Grained LLM-Generated Text Detection via Role Recognition and Involvement MeasurementZihao Cheng, Li Zhou, Feng Jiang, Benyou Wang 等WWW 2025 · 被引用 20 次
- OpenTuringBench: An Open-Model-based Benchmark and Framework for Machine-Generated Text Detection and AttributionLucio La Cava, Andrea TagarelliEMNLP 2025
- MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection BenchmarkDominik Macko, Róbert Móro, Adaku Uchendu, Jason Samuel Lucas 等EMNLP 2023 · 被引用 25 次
