Automatically Exposing Problems with Neural Dialog Models
Dian Yu, Kenji Sagae
Abstract
Neural dialog models are known to suffer from problems such as generating unsafe and inconsistent responses. Even though these problems are crucial and prevalent, they are mostly manually identified by model designers through interactions. Recently, some research instructs crowdworkers to goad the bots into triggering such problems. However, humans leverage superficial clues such as hate speech, while leaving systematic problems undercover. In this paper, we propose two methods including reinforcement learning to automatically trigger a dialog model into generating problematic responses. We show the effect of our methods in exposing safety and contradiction issues with state-of-the-art dialog models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e9d687a-bb22-4b51-a2c1-c9c546f70d90Cited by top-tier papers4
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai et al.EMNLP 2022 · 239 citations
- Fundamental Limitations of Alignment in Large Language ModelsYotam Wolf, Noam Wies, Oshri Avnery, Yoav Levine et al.ICML 2024 · 186 citations
- Reproducibility in Computational Linguistics: Is Source Code Enough?Mohammad Arvan, Luís Pina, Natalie PardeEMNLP 2022 · 12 citations
- Legilimens: Practical and Unified Content Moderation for Large Language Model ServicesJialin Wu, Jiangyi Deng, Shengyuan Pang, Yanjiao Chen et al.CCS 2024 · 5 citations
Builds on12
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung et al.ICLR 2020 · 1,166 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- Beyond Goldfish Memory: Long-Term Open-Domain ConversationJing Xu, Arthur Szlam, Jason WestonACL 2022 · 329 citations
Related papers
- Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-DefenseGuobin Shen, Dongcheng Zhao, Haibo Tong, Jindong Li et al.ICLR 2026 · 4 citations
- More Sail than Ballast: Addressing Harmful Knowledge Leakage in the Expansive Reasoning Space of LRMsQibing Ren, Xinhao Song, Ke Fan, Lijun Li et al.ICML 2026
- Negative Training for Neural Dialogue Response GenerationTianxing He, James R. GlassACL 2020 · 52 citations
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji et al.ICLR 2024 · 656 citations
- SafeConv: Explaining and Correcting Conversational Unsafe BehaviorMian Zhang, Lifeng Jin, Linfeng Song, Haitao Mi et al.ACL 2023 · 5 citations
