Learning From Dictionary: Enhancing Robustness of Machine-Generated Text Detection in Zero-Shot Language via Adversarial Training
Yuanfan Li, Qi Zhou, Zexuan Xie
Abstract
Machine-generated text (MGT) detection is critical for safeguarding online content integrity and preventing the spread of misleading information. Although existing detectors achieve high accuracy in monolingual settings, they exhibit severe performance degradation on zero-shot languages and are vulnerable to adversarial attacks. To tackle these challenges, we propose a robust adversarial training framework named Translation-based Attacker Strengthens MulTilingual DefEnder (TASTE). TASTE comprises two core components: an attacker that performs code-switching by querying translation dictionaries to generate adversarial examples, and a detector trained to resist these attacks while generalizing to unseen languages. We further introduce a novel Language-Agnostic Adversarial Loss (LAAL), which encourages the detector to learn language-invariant feature representations and thus enhances zero-shot detection performance and robustness against unseen attacks. Additionally, the attacker and detector are synchronously updated, enabling continuous improvement of defensive capabilities. Experimental results on 9 languages and 8 attack types show that our TASTE surpasses 8 SOTA detectors, improving the average F1 score by 0.064 and reducing the average Attack Success Rate (ASR) by 3.8%. Our framework offers a promising approach for building robust, multilingual MGT detectors with strong generalization to real-world adversarial scenarios. Our codes are available in https://github.com/Liyuuuu111/MGT-Eval , and our datasets and pretrained checkpoint are available in https://drive.google.com/ file/d/1w1hbdiZMS_JzPntVMWM3qrTQ4KxJf-t6 . REPRODUCIBILITY STATEMENT Code and artifacts. We commit to releasing: (1) training/evaluation code; (2) datasets used in this paper; (3) trained detector checkpoints; and (4) the translation-dictionary resources we used or scripts to construct them from public sources (with licenses). Data and splits. We describe all datasets, languages, and splits used for training/validation/test in Appendix A and provide scripts to regenerate them deterministically from public releases. Training details. We enumerate all hyperparameters (optimizer, learning rates for detector/surrogate, batch sizes, epochs, fp16, gradient clipping), schedules (e.g., GRL weight schedule), and attack-strength curriculum (initial tokens, maximum ratio, increment per step) in Appendix A.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 647ffcf9-9437-47b7-a290-2700229366b7Builds on18
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning et al.ICML 2023 · 988 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
Related papers
- Iron Sharpens Iron: Defending Against Attacks in Machine-Generated Text Detection with Adversarial TrainingYuanfan Li, Zhaohan Zhang, Chengzhengxu Li, Chao Shen et al.ACL 2025 · 10 citations
- OSTAR: Optimized Statistical Text-classifier with Adversarial ResistanceYuhan Yao, Feifei Kou, Lei Shi, Xiao Yang et al.NeurIPS 2025
- Deepfake Text Detection: Limitations and OpportunitiesJiameng Pu, Zain Sarwar, Sifat Muhammad Abdullah, Abdullah Rehman et al.S&P 2023
- Stumbling Blocks: Stress Testing the Robustness of Machine-Generated Text Detectors Under AttacksYichen Wang, Shangbin Feng, Abe Bohan Hou, Xiao Pu et al.ACL 2024
- Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated TextYize Cheng, Vinu Sankar Sadasivan, Mehrdad Saberi, Shoumik Saha et al.NeurIPS 2025 · 28 citations
