Contextual Gaps in Machine Learning for Mental Illness Prediction: The Case of Diagnostic Disclosures
Stevie Chancellor, Jessica L. Feuston, Jayhyun Chang
Abstract
Getting training data for machine learning (ML) prediction of mental illness on social media data is labor intensive. To work around this, ML teams will extrapolate proxy signals, or alternative signs from data to evaluate illness status and create training datasets. However, these signals' validity has not been determined, whether signals align with important contextual factors, and how proxy quality impacts downstream model integrity. We use ML and qualitative methods to evaluate whether a popular proxy signal, diagnostic self-disclosure, produces a conceptually sound ML model of mental illness. Our findings identify major conceptual errors only seen through a qualitative investigation -- training data built from diagnostic disclosures encodes a narrow vision of diagnosis experiences that propagates into paradoxes in the downstream ML model. This gap is obscured by strong performance of the ML classifier (F1 = 0.91). We discuss the implications of conceptual gaps in creating training data for human-centered models, and make suggestions for improving research methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 256dcac1-bdd1-42ad-889a-44cb9ece000bCited by top-tier papers1
Ask how each one uses itRelated papers
- A Simple and Flexible Modeling for Mental Disorder Detection by Learning from Clinical QuestionnairesHoyun Song, Jisu Shin, Huije Lee, Jong C. ParkACL 2023
- Improving the Generalizability of Depression Detection by Leveraging Clinical QuestionnairesThong Nguyen, Andrew Yates, Ayah Zirikly, Bart Desmet et al.ACL 2022 · 66 citations
- DRMD: Explainable Depression Detection Based on Metaphorical Conceptual MappingDongyu Zhang, Wanqiu Liao, Weichen Hu, Hongfei LinWWW 2026
- Semantic Gap in Predicting Mental Wellbeing through Passive SensingVedant Das Swain, Victor Chen, Shrija Mishra, Stephen M. Mattingly et al.CHI 2022 · 36 citations
- What is Stigma Attributed to? A Theory-Grounded, Expert-Annotated Interview Corpus for Demystifying Mental-Health StigmaHan Meng, Yancan Chen, Yunan Li, Yitian Yang et al.ACL 2025 · 5 citations
