Trigger Warning Assignment as a Multi-Label Document Classification Problem
Matti Wiegmann, Magdalena Wolska, Christopher Schröder, Ole Borchardt, Benno Stein, Martin Potthast
Abstract
A trigger warning is used to warn people about potentially disturbing content. We introduce trigger warning assignment as a multi-label classification task, create the Webis Trigger Warning Corpus 2022, and with it the first dataset of 1 million fanfiction works from Archive of our Own with up to 36 different warnings per document. To provide a reliable catalog of trigger warnings, we organized 41 million of free-form tags assigned by fanfiction authors into the first comprehensive taxonomy of trigger warnings by mapping them to the 36 institutionally recommended warnings. To determine the best operationalization of trigger warnings, we explore state-of-the-art multi-label models, examining the trade-off between assigning coarse-and fine-grained warnings, open-and closed-set classification, document length, and label confidence. Our models achieve micro-F 1 scores of about 0.5, which reveals the difficulty of the task. Tailored representations, long input sequences, and a higher recall on rare warnings would help. 1,2
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2bc84bd-f17b-4597-bb59-d74c5e414c33Builds on1
Related papers
- Towards Understanding Unsafe Video GenerationYan Pang, Aiping Xiong, Yang Zhang, Tianhao WangNDSS 2025
- T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image ModelChenyu Zhang, Tairen Zhang, Lanjun Wang, Ruidong Chen et al.AAAI 2026 · 3 citations
- Bridging the Copyright Gap: Do Large Vision-Language Models Recognize and Respect Copyrighted Content?Naen Xu, Jinghuai Zhang, Changjiang Li, Hengyu An et al.AAAI 2026 · 6 citations
- ArtEmis: Affective Language for Visual ArtPanos Achlioptas, Maks Ovsjanikov, Kilichbek Haydarov, Mohamed Elhoseiny et al.CVPR 2021
- Ruddit: Norms of Offensiveness for English Reddit CommentsRishav Hada, Sohi Sudhir, Pushkar Mishra, Helen Yannakoudakis et al.ACL 2021
