Donate or Create? Comparing Data Collection Strategies for Emotion-labeled Multimodal Social Media Posts
Christopher Bagdon, Aidan Combs, Carina Silberer, Roman Klinger
摘要
Accurate modeling of subjective phenomena such as emotion expression requires data annotated with authors' intentions. Commonly such data is collected by asking study participants to donate and label genuine content produced in the real world, or create content fitting particular labels during the study. Asking participants to create content is often simpler to implement and presents fewer risks to participant privacy than data donation. However, it is unclear if and how study-created content may differ from genuine content, and how differences may impact models. We collect study-created and genuine multimodal social media posts labeled for emotion and compare them on several dimensions, including model performance. We find that compared to genuine posts, study-created posts are longer, rely more on their text and less on their images for emotion expression, and focus more on emotion-prototypical events. The samples of participants willing to donate versus create posts are demographically different. Study-created data is valuable to train models that generalize well to genuine data, but realistic effectiveness estimates require genuine data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 被引用 51 次
- Out-of-Distribution Generalization in Natural Language Processing: Past, Present, and FutureLinyi Yang, Yaoxian Song, Xuan Ren, Chenyang Lyu 等EMNLP 2023 · 被引用 12 次
- iSarcasm: A Dataset of Intended SarcasmSilviu Oprea, Walid MagdyACL 2020 · 被引用 2 次
相关 Paper
- Can Third Parties Read Our Emotions?Jiayi Li, Yingfan Zhou, Pranav Narayanan Venkit, Halima Binte Islam 等ACL 2025 · 被引用 5 次
- Contextual Gaps in Machine Learning for Mental Illness Prediction: The Case of Diagnostic DisclosuresStevie Chancellor, Jessica L. Feuston, Jayhyun ChangCSCW 2023 · 被引用 6 次
- Detecting Users' Emotional States during Passive Social Media UseChristoph Gebhardt, Andreas Brombach, Tiffany Luong, Otmar Hilliges 等UbiComp 2024 · 被引用 8 次
- "Finsta gets all my bad pictures": Instagram Users' Self-Presentation Across Finsta and Rinsta AccountsXiaoyun Huang, Jessica VitakCSCW 2022 · 被引用 39 次
- Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text PerceptionsMatthias Orlikowski, Jiaxin Pei, Paul Röttger, Philipp Cimiano 等ACL 2025 · 被引用 34 次
