SafeText: A Benchmark for Exploring Physical Safety in Language Models
Sharon Levy, Emily Allaway, Melanie Subbiah, Lydia B. Chilton, Desmond Patton, Kathleen R. McKeown, William Yang Wang
Abstract
Understanding what constitutes safe text is an important issue in natural language processing and can often prevent the deployment of models deemed harmful and unsafe. One such type of safety that has been scarcely studied is commonsense physical safety, i.e. text that is not explicitly violent and requires additional commonsense knowledge to comprehend that it leads to physical harm. We create the first benchmark dataset, SAFETEXT, comprising real-life scenarios with paired safe and physically unsafe pieces of advice. We utilize SAFE-TEXT to empirically study commonsense physical safety across various models designed for text generation and commonsense reasoning tasks. We find that state-of-the-art large language models are susceptible to the generation of unsafe text and have difficulty rejecting unsafe advice. As a result, we argue for further studies of safety and the assessment of commonsense physical safety in models before release.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ffe77761-8ce3-45f1-b71d-2d3e38db1c61Cited by top-tier papers9
- Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow InstructionsFederico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger et al.ICLR 2024 · 373 citations
- ROBBIE: Robust Bias Evaluation of Large Generative Language ModelsDavid Esiobu, Xiaoqing Ellen Tan, Saghar Hosseini, Megan Ung et al.EMNLP 2023 · 16 citations
- Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy OptimizationXiyue Peng, Hengquan Guo, Jiawei Zhang, Dongqing Zou et al.NeurIPS 2025 · 9 citations
- GuardBench: A Large-Scale Benchmark for Guardrail ModelsElias Bassani, Ignacio SanchezEMNLP 2024 · 7 citations
- CRoW: Benchmarking Commonsense Reasoning in Real-World TasksMete Ismayilzada, Debjit Paul, Syrielle Montariol, Mor Geva et al.EMNLP 2023 · 3 citations
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch et al.ICLR 2021 · 878 citations
- (Comet-) Atomic 2020: On Symbolic and Neural Commonsense Knowledge GraphsJena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da et al.AAAI 2021 · 458 citations
Related papers
- Subtle Risks, Critical Failures: A Framework for Diagnosing Physical Safety of LLMs for Embodied Decision MakingYejin Son, Minseo Kim, Sungwoong Kim, Seungju Han et al.EMNLP 2025 · 7 citations
- A Benchmark for Semantic Sensitive Information in LLMs OutputsQingjie Zhang, Han Qiu, Di Wang, Yiming Li et al.ICLR 2025
- Multimodal Situational SafetyKaiwen Zhou, Chengzhi Liu, Xuandong Zhao, Anderson Compalas et al.ICLR 2025
- Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMsWenxuan Wang, Xiaoyuan Liu, Kuiyi Gao, Jen-tse Huang et al.ACL 2025
- SafetyBench: Evaluating the Safety of Large Language ModelsZhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun et al.ACL 2024
