Stanceosaurus: Classifying Stance Towards Multicultural Misinformation
Jonathan Zheng, Ashutosh Baheti, Tarek Naous, Wei Xu, Alan Ritter
摘要
We present Stanceosaurus, a new corpus of 28,033 tweets in English, Hindi, and Arabic annotated with stance towards 251 misinformation claims. As far as we are aware, it is the largest corpus annotated with stance towards misinformation claims. The claims in Stanceosaurus originate from 15 fact-checking sources that cover diverse geographical regions and cultures. Unlike existing stance datasets, we introduce a more fine-grained 5class labeling strategy with additional subcategories to distinguish implicit stance. Pretrained transformer-based stance classifiers that are fine-tuned on our corpus show good generalization on unseen claims and regional claims from countries outside the training data. Cross-lingual experiments demonstrate Stanceosaurus' capability of training multilingual models, achieving 53.1 F1 on Hindi and 50.4 F1 on Arabic without any targetlanguage fine-tuning. Finally, we show how a domain adaptation method can be used to improve performance on Stanceosaurus using additional RumourEval-2019 data. We make Stanceosaurus publicly available to the research community and hope it will encourage further work on misinformation identification across languages and cultures. 1 Dataset Target Number/Range of Topics SemEval-2016 (Mohammad et al., 2016) Subject 6 political topics (e.g., atheism, feminist movement) SRQ (Villa-Cox et al., 2020) Subject 4 political topics & events (e.g., general terms, student marches) Catalonia (Zotova et al., 2020) Subject 1 topic (i.e., Catalonia independence) COVID (Glandt et al., 2021) Subject 4 topic related to Covid-19 (e.g., stay at home orders) Multi-target (Sobhani et al., 2017) Entity 3 pairs of candidates in 2016 US election WTWT (Conforti et al., 2020) Event 5 merger and acquisition events RumourEval (Gorrell et al., 2019) Tweet 8 news events + rumors about natural disasters Rumor-has-it (Qazvinian et al., 2011) Claim 5 rumors (e.g., Sarah Palin getting divorced?) CovidLies (Hossain et al., 2020) Claim 86 pieces of COVID-19 misinformation Stanceosaurus (this work) Claim 251 claims over a diverse set of global and regional topics Table 1: Summary of Twitter stance classification datasets. Stanceosaurus covers more claims from a broader range of topics and geographical regions than prior Twitter stance datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Towards Understanding Factual Knowledge of Large Language ModelsXuming Hu, Junzhe Chen, Xiaochuan Li, Yufei Guo 等ICLR 2024 · 被引用 21 次
- On the Risk of Evidence Pollution for Malicious Social Text Detection in the Era of LLMsHerun Wan, Minnan Luo, Zhixiong Su, Guang Dai 等ACL 2025 · 被引用 5 次
- Enabling Contextual Soft Moderation on Social Media through Contrastive Textual DeviationPujan Paudel, Mohammad Hammas Saeed, Rebecca Auger, Chris Wells 等USENIX Security 2024 · 被引用 3 次
- How to Protect Yourself from 5G Radiation? Investigating LLM Responses to Implicit MisinformationRuohao Guo, Wei Xu, Alan RitterEMNLP 2025
- To Lie or Not to Lie? Investigating The Biased Spread of Global Lies by LLMsZohaib Khan, Mustafa Dogan, Ifeoma Okoh, Pouya Sadeghi 等ACL 2026
它引用的顶会 Paper6
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Coupled Hierarchical Transformer for Stance-Aware Rumor Verification in Social Media ConversationsJianfei Yu, Jing Jiang, Ling Min Serena Khoo, Hai Leong Chieu 等EMNLP 2020 · 被引用 49 次
- A Weakly Supervised Propagation Model for Rumor Verification and Stance Detection with Multiple Instance LearningRuichao Yang, Jing Ma, Hongzhan Lin, Wei GaoSIGIR 2022 · 被引用 39 次
- Cross-Domain Label-Adaptive Stance DetectionMomchil Hardalov, Arnav Arora, Preslav Nakov, Isabelle AugensteinEMNLP 2021 · 被引用 3 次
相关 Paper
- Stance Detection in COVID-19 TweetsKyle Glandt, Sarthak Khanal, Yingjie Li, Doina Caragea 等ACL 2021
- COVID-19 Vaccine Misinformation in Middle Income CountriesJongin Kim, Byeo Bak, Aditya Agrawal, Jiaxi Wu 等EMNLP 2023 · 被引用 3 次
- Multilingual Previously Fact-Checked Claim RetrievalMatús Pikuliak, Ivan Srba, Róbert Móro, Timo Hromadka 等EMNLP 2023 · 被引用 9 次
- The Surprising Performance of Simple Baselines for Misinformation DetectionKellin Pelrine, Jacob Danovitch, Reihaneh RabbanyWWW 2021 · 被引用 79 次
- From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP ModelsShangbin Feng, Chan Young Park, Yuhan Liu, Yulia TsvetkovACL 2023 · 被引用 117 次
