SQuARe: A Large-Scale Dataset of Sensitive Questions and Acceptable Responses Created through Human-Machine Collaboration
Hwaran Lee, Seokhee Hong, Joonsuk Park, Takyoung Kim, Meeyoung Cha, Yejin Choi, Byoung Pil Kim, Gunhee Kim, Eun-Ju Lee, Yong Lim, Alice Oh, Sangchul Park, Jung-Woo Ha
Abstract
The potential social harms that large language models pose, such as generating offensive content and reinforcing biases, are steeply rising. Existing works focus on coping with this concern while interacting with ill-intentioned users, such as those who explicitly make hate speech or elicit harmful responses. However, discussions on sensitive issues can become toxic even if the users are well-intentioned. For safer models in such scenarios, we present the Sensitive Questions and Acceptable Response (SQUARE) dataset, a large-scale Korean dataset of 49k sensitive questions with 42k acceptable and 46k non-acceptable responses. The dataset was constructed leveraging HyperCLOVA in a human-in-the-loop manner based on real news headlines. Experiments show that acceptable response generation significantly improves for HyperCLOVA and GPT-3, demonstrating the efficacy of this dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80ad0f07-a63d-4811-97fb-2aa3aeb2a7d0Cited by top-tier papers3
- CAGE: A Framework for Culturally Adaptive Red-Teaming Benchmark GenerationChaeyun Kim, YongTaek Lim, Kihyun Kim, Junghwan Kim et al.ICLR 2026 · 2 citations
- Delving into Multilingual Ethical Bias: The MSQAD with Statistical Hypothesis Tests for Large Language ModelsSeunguk Yu, Juhwan Choi, YoungBin KimACL 2025 · 2 citations
- Social Bias in Multilingual Language Models: A SurveyLance Calvin Lim Gamboa, Yue Feng, Mark G. LeeEMNLP 2025
Builds on17
- Cross-Task Generalization via Natural Language Crowdsourcing InstructionsSwaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh HajishirziACL 2022 · 887 citations
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch et al.ICLR 2021 · 878 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Process for Adapting Language Models to Society (PALMS) with Values-Targeted DatasetsIrene Solaiman, Christy DennisonNeurIPS 2021 · 276 citations
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai et al.EMNLP 2022 · 239 citations
Related papers
- ProsocialDialog: A Prosocial Backbone for Conversational AgentsHyunwoo Kim, Youngjae Yu, Liwei Jiang, Ximing Lu et al.EMNLP 2022 · 46 citations
- PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human PreferenceJiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen et al.ACL 2025
- A Benchmark for Semantic Sensitive Information in LLMs OutputsQingjie Zhang, Han Qiu, Di Wang, Yiming Li et al.ICLR 2025
- What Changes Can Large-scale Language Models Bring? Intensive Study on HyperCLOVA: Billions-scale Korean Generative Pretrained TransformersBoseop Kim, HyoungSeok Kim, Sang-Woo Lee, Gichang Lee et al.EMNLP 2021 · 11 citations
- CoBia: Constructed Conversations Can Trigger Otherwise Concealed Societal Biases in LLMsNafiseh Nikeghbal, Amir Hossein Kargaran, Jana DiesnerEMNLP 2025
