MythTriage: Scalable Detection of Opioid Use Disorder Myths on a Video-Sharing Platform
Hayoung Jung, Shravika Mittal, Ananya Aatreya, Navreet Kaur, Munmun De Choudhury, Tanushree Mitra
Abstract
Understanding the prevalence of misinformation in health topics online can inform public health policies and interventions. However, measuring such misinformation at scale remains a challenge, particularly for high-stakes but understudied topics like opioid-use disorder (OUD)-a leading cause of death in the U.S. We present the first large-scale study of OUDrelated myths on YouTube, a widely-used platform for health information. With clinical experts, we validate 8 pervasive myths and release an expert-labeled video dataset. To scale labeling, we introduce MYTHTRIAGE, an efficient triage pipeline that uses a lightweight model for routine cases and defers harder ones to a high-performing, but costlier, large language model (LLM). MYTHTRIAGE achieves up to 0.86 macro F1-score while estimated to reduce annotation time and financial cost by over 76% compared to experts and full LLM labeling. We analyze 2.9K search results and 343K recommendations, uncovering how myths persist on YouTube and offering actionable insights for public health and platform moderation. 1 Warning: Some content of this paper, included to contextualize our data, is misleading. 1. Data Collection 2. Data Labeling for OUD-Related Myths 1.1: Curating Opioid Topics and Queries Google Trends YouTube Autocomplete Fentanyl 8 Topics Kratom … 73 Queries fentanyl overdose fentanyl … 1.2: Collecting Data on YouTube Collecting Search Results Queries Collect Top-10 Results OUD Search Dataset (2.9K results) Repeat x4 Search Filter Collecting Recommendation Results (Recs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b11f95ff-10d5-40d3-9946-9e8b93feb27cCited by top-tier papers1
Ask how each one uses itBuilds on12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Misinformation in Video Search Platforms: An Audit Study on YouTubeEslam Hussein, Prerna Juneja, Tanushree MitraCSCW 2020 · 233 citations
- MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical ReasoningShuyue Stella Li, Vidhisha Balachandran, Shangbin Feng, Jonathan Ilgen et al.NeurIPS 2024 · 215 citations
- Large Language Models for Data Annotation and Synthesis: A SurveyZhen Tan, Dawei Li, Song Wang, Alimohammad Beigi et al.EMNLP 2024 · 119 citations
- The Typing Cure: Experiences with Large Language Model Chatbots for Mental Health SupportInhwa Song, Sachin R. Pendse, Neha Kumar, Munmun De ChoudhuryCSCW 2025 · 41 citations
Related papers
- From Knowledge to Inference: Formalizing Specialized Public Health Reasoning on GlobalHealthAtlasZhaokun Yan, Shan Xu, Wuzheng Dong, Zhaohan Liu et al.ICML 2026
- Education, Personal Experiences, and Advocacy: Examining Drug-Addiction Videos on YouTubeShuo Niu, Katherine G. McKim, Kathleen Palm ReedCSCW 2022 · 11 citations
- MentalSeek-Dx: Towards Progressive Hypothetico-Deductive Reasoning for Real-world Psychiatric DiagnosisXiao Sun, Yuming Yang, Xinyi Jiang, Yu Tian et al.ACL 2026 · 1 citation
- Not all Fake News is Written: A Dataset and Analysis of Misleading Video HeadlinesYoo Yeon Sung, Jordan L. Boyd-Graber, Naeemul HassanEMNLP 2023 · 1 citation
- Missci: Reconstructing Fallacies in Misrepresented ScienceMax Glockner, Yufang Hou, Preslav Nakov, Iryna GurevychACL 2024 · 1 citation
