OSTAR: Optimized Statistical Text-classifier with Adversarial Resistance
Yuhan Yao, Feifei Kou, Lei Shi, Xiao Yang, Zhongbao Zhang, Suguo Zhu, Jiwei Zhang, Lirong Qiu, Haisheng Li
Abstract
The advancements in generative models and the real-world attack of machine-generated text(MGT) create a demand for more robust detection methods. The existing MGT detection methods for adversarial environments primarily consist of manually designed statistical-based methods and fine-tuned classifier-based approaches. Statistical-based methods extract intrinsic features but suffer from rigid decision boundaries vulnerable to adaptive attacks, while fine-tuned classifiers achieve outstanding performance at the cost of overfitting to superficial textual feature. We argue that the key to detection in current adversarial environments lies in how to extract intrinsic invariant features and ensure that the classifier possesses dynamic adaptability. In that case, we propose OSTAR , a novel MGT detection framework designed for adversarial environments which composed of a statistical enhanced classifier and a Multi-Faceted Contrastive Learning(MFCL). In the classifier aspect, our Multi-Dimensional Statistical Profiling (MDSP) module extracts intrinsic difference between human and machine texts, complementing classifiers with useful stable features. In the model optimization aspect, the MFCL strategy enhances robustness by contrasting feature variations before and after text attacks, jointly optimizing statistical feature mapping and baseline pre-trained models. Experimental results on three public datasets under various adversarial scenarios demonstrate that our framework outperforms existing MGT detection methods, achieving state-of-the-art performance and robust against attacks.The code is available at https://github.com/BUPT-SN/OSTAR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 74193e6b-cc70-460d-bdf6-830bd05a6c69Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning et al.ICML 2023 · 988 citations
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseKalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting et al.NeurIPS 2023 · 657 citations
- RADAR: Robust AI-Text Detection via Adversarial LearningXiaomeng Hu, Pin-Yu Chen, Tsung-Yi HoNeurIPS 2023 · 315 citations
Related papers
- Learning From Dictionary: Enhancing Robustness of Machine-Generated Text Detection in Zero-Shot Language via Adversarial TrainingYuanfan Li, Qi Zhou, Zexuan XieICLR 2026
- MGTBench: Benchmarking Machine-Generated Text DetectionXinlei He, Xinyue Shen, Zeyuan Chen, Michael Backes et al.CCS 2024 · 30 citations
- Iron Sharpens Iron: Defending Against Attacks in Machine-Generated Text Detection with Adversarial TrainingYuanfan Li, Zhaohan Zhang, Chengzhengxu Li, Chao Shen et al.ACL 2025 · 10 citations
- CoCo: Coherence-Enhanced Machine-Generated Text Detection Under Low Resource With Contrastive LearningXiaoming Liu, Zhaohan Zhang, Yichen Wang, Hang Pu et al.EMNLP 2023 · 16 citations
- DeTeCtive: Detecting AI-generated Text via Multi-Level Contrastive LearningXun Guo, Yongxin He, Shan Zhang, Ting Zhang et al.NeurIPS 2024 · 100 citations
