Stanceosaurus: Classifying Stance Towards Multicultural Misinformation
Jonathan Zheng, Ashutosh Baheti, Tarek Naous, Wei Xu, Alan Ritter
Abstract
We present Stanceosaurus, a new corpus of 28,033 tweets in English, Hindi, and Arabic annotated with stance towards 251 misinformation claims. As far as we are aware, it is the largest corpus annotated with stance towards misinformation claims. The claims in Stanceosaurus originate from 15 fact-checking sources that cover diverse geographical regions and cultures. Unlike existing stance datasets, we introduce a more fine-grained 5class labeling strategy with additional subcategories to distinguish implicit stance. Pretrained transformer-based stance classifiers that are fine-tuned on our corpus show good generalization on unseen claims and regional claims from countries outside the training data. Cross-lingual experiments demonstrate Stanceosaurus' capability of training multilingual models, achieving 53.1 F1 on Hindi and 50.4 F1 on Arabic without any targetlanguage fine-tuning. Finally, we show how a domain adaptation method can be used to improve performance on Stanceosaurus using additional RumourEval-2019 data. We make Stanceosaurus publicly available to the research community and hope it will encourage further work on misinformation identification across languages and cultures. 1 Dataset Target Number/Range of Topics SemEval-2016 (Mohammad et al., 2016) Subject 6 political topics (e.g., atheism, feminist movement) SRQ (Villa-Cox et al., 2020) Subject 4 political topics & events (e.g., general terms, student marches) Catalonia (Zotova et al., 2020) Subject 1 topic (i.e., Catalonia independence) COVID (Glandt et al., 2021) Subject 4 topic related to Covid-19 (e.g., stay at home orders) Multi-target (Sobhani et al., 2017) Entity 3 pairs of candidates in 2016 US election WTWT (Conforti et al., 2020) Event 5 merger and acquisition events RumourEval (Gorrell et al., 2019) Tweet 8 news events + rumors about natural disasters Rumor-has-it (Qazvinian et al., 2011) Claim 5 rumors (e.g., Sarah Palin getting divorced?) CovidLies (Hossain et al., 2020) Claim 86 pieces of COVID-19 misinformation Stanceosaurus (this work) Claim 251 claims over a diverse set of global and regional topics Table 1: Summary of Twitter stance classification datasets. Stanceosaurus covers more claims from a broader range of topics and geographical regions than prior Twitter stance datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3c20987-e313-436c-b96a-bd90e10774c1Cited by top-tier papers5
- Towards Understanding Factual Knowledge of Large Language ModelsXuming Hu, Junzhe Chen, Xiaochuan Li, Yufei Guo et al.ICLR 2024 · 21 citations
- On the Risk of Evidence Pollution for Malicious Social Text Detection in the Era of LLMsHerun Wan, Minnan Luo, Zhixiong Su, Guang Dai et al.ACL 2025 · 5 citations
- Enabling Contextual Soft Moderation on Social Media through Contrastive Textual DeviationPujan Paudel, Mohammad Hammas Saeed, Rebecca Auger, Chris Wells et al.USENIX Security 2024 · 3 citations
- How to Protect Yourself from 5G Radiation? Investigating LLM Responses to Implicit MisinformationRuohao Guo, Wei Xu, Alan RitterEMNLP 2025
- To Lie or Not to Lie? Investigating The Biased Spread of Global Lies by LLMsZohaib Khan, Mustafa Dogan, Ifeoma Okoh, Pouya Sadeghi et al.ACL 2026
Builds on6
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Coupled Hierarchical Transformer for Stance-Aware Rumor Verification in Social Media ConversationsJianfei Yu, Jing Jiang, Ling Min Serena Khoo, Hai Leong Chieu et al.EMNLP 2020 · 49 citations
- A Weakly Supervised Propagation Model for Rumor Verification and Stance Detection with Multiple Instance LearningRuichao Yang, Jing Ma, Hongzhan Lin, Wei GaoSIGIR 2022 · 39 citations
- Cross-Domain Label-Adaptive Stance DetectionMomchil Hardalov, Arnav Arora, Preslav Nakov, Isabelle AugensteinEMNLP 2021 · 3 citations
Related papers
- Stance Detection in COVID-19 TweetsKyle Glandt, Sarthak Khanal, Yingjie Li, Doina Caragea et al.ACL 2021
- COVID-19 Vaccine Misinformation in Middle Income CountriesJongin Kim, Byeo Bak, Aditya Agrawal, Jiaxi Wu et al.EMNLP 2023 · 3 citations
- Multilingual Previously Fact-Checked Claim RetrievalMatús Pikuliak, Ivan Srba, Róbert Móro, Timo Hromadka et al.EMNLP 2023 · 9 citations
- The Surprising Performance of Simple Baselines for Misinformation DetectionKellin Pelrine, Jacob Danovitch, Reihaneh RabbanyWWW 2021 · 79 citations
- From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP ModelsShangbin Feng, Chan Young Park, Yuhan Liu, Yulia TsvetkovACL 2023 · 117 citations
