Reinforcement Learning-Guided Adaptive Tuning for Out-of-Distribution Harmful Text Detection
Mengyu Xiang, Tinghao Chen, Boxu Han, Qiudan Li, Shu Wu, Daniel Dajun Zeng
Abstract
As social media grows, harmful information spreads rapidly across platforms and evolves over time, showing cross-platform and cross-temporal variations. Existing methods rely on fixed model parameters during training, which fail to handle substantial semantic discrepancies, leading to Out-Of-Distribution (OOD) problems. While test-time tuning enables dynamic parameter adjustment, it may lead to excessive adaptation to individual samples. The key challenge is how to adapt to semantic variations during testing while preventing overfitting from continuous tuning. To tackle this issue, this paper proposes RLAT, a reinforcement learning (RL)–guided adaptive tuning method for harmful text detection. First, a tuning joint optimization module is designed to update parameters and adapt to semantic variations during testing. It tunes the model by optimizing consistency loss and applying word-level attention constraints to reduce over-reliance on local words and learn a more robust global representation. Then, to mitigate overfitting caused by continuous tuning, a RL–guided adaptive decision model is introduced to direct the tuning process. It reduces the influence of local samples by selecting data and controlling parameter updates, thereby improving overall test performance. Experimental results show that the RLAT outperforms state-of-the-art baselines in cross-platform and cross-temporal scenarios across multiple public datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Is Sarcasm Detection a Step-by-Step Reasoning Process in Large Language Models?Ben Yao, Yazhou Zhang, Qiuchi Li, Jing QinAAAI 2025 · 29 citations
- Facilitating Fine-grained Detection of Chinese Toxic Language: Hierarchical Taxonomy, Resources, and BenchmarksJunyu Lu, Bo Xu, Xiaokun Zhang, Changrong Min et al.ACL 2023 · 25 citations
- Sheep's Skin, Wolf's Deeds: Are LLMs Ready for Metaphorical Implicit Hate Speech?Jingjie Zeng, Liang Yang, Zekun Wang, Yuanyuan Sun et al.ACL 2025 · 4 citations
- Attention-guided Self-reflection for Zero-shot Hallucination Detection in Large Language ModelsQiang Liu, Xinlong Chen, Yue Ding, Bowen Song et al.EMNLP 2025 · 2 citations
Related papers
- Shedding the Facades, Connecting the Domains: Detecting Shifting Multimodal Hate Video with Test-Time AdaptationJiao Li, Jian Lang, Xikai Tang, Wenzheng Shu et al.AAAI 2026
- Towards Policy-Adaptive Image Guardrail: Benchmark and MethodCaiyong Piao, Zhiyuan Yan, Haoming Xu, Yunzhen Zhao et al.CVPR 2026 · 7 citations
- Selective Test-Time Debiasing for CLIP via Reward GatingJaeho Han, Jisoo Yang, Hyeondong Woo, Mingyu Jeon et al.ACL 2026
- ContiGuard: A Framework for Continual Toxicity Detection Against Evolving Evasive PerturbationsHankun Kang, Xin Miao, Jianhao Chen, Jintao Wen et al.WWW 2026
- Boosting ASR Robustness via Test-Time Reinforcement Learning with Audio-Text Semantic RewardsLinghan Fang, Tianxin Xie, Li LiuAAAI 2026 · 1 citation
