DAC: Quantized Optimal Transport Reward-based Reinforcement Learning Approach to Detoxify Query Auto-Completion
Aishwarya Maheswaran, Kaushal Kumar Maurya, Manish Gupta, Maunendra Sankar Desarkar
Abstract
Modern Query Auto-Completion (QAC) systems utilize natural language generation (NLG) using large language models (LLM) to achieve remarkable performance. However, these systems are prone to generating biased and toxic completions due to inherent learning biases. Existing detoxification approaches exhibit two key limitations: (1) They primarily focus on mitigating toxicity for grammatically well-formed long sentences but struggle to adapt to the QAC task, where queries are short and structurally different (include spelling errors, do not follow grammatical rules and have relatively flexible word order). (2) These approaches often view detoxification through a binary lens where all text labeled as toxic is undesirable, and non-toxic is considered desirable. To address these limitations, we propose DAC, an intuitive and efficient reinforcement learning-based model to detoxify QAC. With DAC, we introduce an additional perspective of considering the third query class of addressable toxicity. These queries can encompass implicit toxicity, subjective toxicity, or non-toxic queries containing toxic words. We incorporate this three-class query behavior perspective into the proposed model through quantized optimal transport to learn distinctions and generate truly non-toxic completions. We evaluate toxicity levels in the generated completions by DAC across two real-world QAC datasets (Bing and AOL) using two classifiers: a publicly available generic classifier (Detoxify) and a search query-specific classifier, which we develop (TClassify). We find that DAC consistently outperforms all existing baselines on the Bing dataset and achieves competitive performance on the AOL dataset for query detoxification. % providing high quality and low toxicity. We make the code publicly available.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get d909679a-e475-487a-8b2d-3468ef740d1fCited by top-tier papers2
- DocQAC: Adaptive Trie-Guided Decoding for Effective In-Document Query Auto-CompletionRahul Mehta, Kavin R. V, Indrajit Pal, Tushar Abhishek et al.SIGIR 2026
- Text Detoxification: Data Efficiency, Semantic Preservation and Model GeneralizationJing Yu, Yibo Zhao, Jiapeng Zhu, Wenming Shao et al.EMNLP 2025
Related papers
- From Chaos to Cure: A Prefix Heuristics Guided Model-Agnostic Adaptive Detoxification FrameworkYuhu Shang, Xiang Cheng, Yimeng Ren, Huijia Wu et al.AAAI 2026
- Systematic Rectification of Language Models via Dead-end AnalysisMeng Cao, Mehdi Fatemi, Jackie C. K. Cheung, Samira ShabanianICLR 2023 · 2 citations
- Detoxifying Large Language Models via Autoregressive Reward Guided Representation EditingYisong Xiao, Aishan Liu, Siyuan Liang, Zonghao Ying et al.NeurIPS 2025 · 12 citations
- Test-Time Detoxification without Training or Learning AnythingBaturay Saglam, Dionysios KalogeriasICML 2026 · 2 citations
- Unveiling the Implicit Toxicity in Large Language ModelsJiaxin Wen, Pei Ke, Hao Sun, Zhexin Zhang et al.EMNLP 2023 · 21 citations
