MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection Benchmark
Dominik Macko, Róbert Móro, Adaku Uchendu, Jason Samuel Lucas, Michiharu Yamashita, Matús Pikuliak, Ivan Srba, Thai Le, Dongwon Lee, Jakub Simko, Mária Bieliková
Abstract
There is a lack of research into capabilities of recent LLMs to generate convincing text in languages other than English and into performance of detectors of machine-generated text in multilingual settings. This is also reflected in the available benchmarks which lack authentic texts in languages other than English and predominantly cover older generators. To fill this gap, we introduce MULTITuDE 1 , a novel benchmarking dataset for multilingual machine-generated text detection comprising of 74,081 authentic and machine-generated texts in 11 languages (ar, ca, cs, de, en, es, nl, pt, ru, uk, and zh) generated by 8 multilingual LLMs. Using this benchmark, we compare the performance of zero-shot (statistical and black-box) and fine-tuned detectors. Considering the multilinguality, we evaluate 1) how these detectors generalize to unseen languages (linguistically similar as well as dissimilar) and unseen LLMs and 2) whether the detectors improve their performance when trained on multiple languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b949c65e-eca9-44f5-9f51-dda7a27e4ad3Cited by top-tier papers17
- Beyond Binary: Towards Fine-Grained LLM-Generated Text Detection via Role Recognition and Involvement MeasurementZihao Cheng, Li Zhou, Feng Jiang, Benyou Wang et al.WWW 2025 · 20 citations
- On the Detectability of ChatGPT Content: Benchmarking, Methodology, and Evaluation through the Lens of Academic WritingZeyan Liu, Zijun Yao, Fengjun Li, Bo LuoCCS 2024 · 18 citations
- RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text DetectorsLiam Dugan, Alyssa Hwang, Filip Trhlík, Andrew Zhu et al.ACL 2024 · 18 citations
- Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation GenerationAneta Zugecova, Dominik Macko, Ivan Srba, Róbert Móro et al.ACL 2025 · 18 citations
- MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media TextsDominik Macko, Jakub Kopal, Róbert Móro, Ivan SrbaACL 2025 · 15 citations
Builds on6
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning et al.ICML 2023 · 988 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Authorship Attribution for Neural Text GenerationAdaku Uchendu, Thai Le, Kai Shu, Dongwon LeeEMNLP 2020 · 110 citations
- Neural Deepfake Detection with Factual Structure of TextWanjun Zhong, Duyu Tang, Zenan Xu, Ruize Wang et al.EMNLP 2020 · 43 citations
Related papers
- M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text DetectionYuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su et al.ACL 2024
- TSM-Bench: Detecting LLM-Generated Text in Real-World Wikipedia Editing PracticesGerrit Quaremba, Elizabeth Black, Denny Vrandecic, Elena SimperlICLR 2026 · 2 citations
- Authorship Attribution in Multilingual Machine-Generated TextsLucio La Cava, Dominik Macko, Róbert Móro, Ivan Srba et al.ACL 2026 · 7 citations
- MAGE: Machine-generated Text Detection in the WildYafu Li, Qintong Li, Leyang Cui, Wei Bi et al.ACL 2024 · 44 citations
- DetectAnyLLM: Towards Generalizable and Robust Detection of Machine-Generated Text Across Domains and ModelsJiachen Fu, Chun-Le Guo, Chongyi LiACM MM 2025 · 4 citations
