Donate or Create? Comparing Data Collection Strategies for Emotion-labeled Multimodal Social Media Posts
Christopher Bagdon, Aidan Combs, Carina Silberer, Roman Klinger
Abstract
Accurate modeling of subjective phenomena such as emotion expression requires data annotated with authors' intentions. Commonly such data is collected by asking study participants to donate and label genuine content produced in the real world, or create content fitting particular labels during the study. Asking participants to create content is often simpler to implement and presents fewer risks to participant privacy than data donation. However, it is unclear if and how study-created content may differ from genuine content, and how differences may impact models. We collect study-created and genuine multimodal social media posts labeled for emotion and compare them on several dimensions, including model performance. We find that compared to genuine posts, study-created posts are longer, rely more on their text and less on their images for emotion expression, and focus more on emotion-prototypical events. The samples of participants willing to donate versus create posts are demographically different. Study-created data is valuable to train models that generalize well to genuine data, but realistic effectiveness estimates require genuine data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 32f85b50-9700-4e00-836a-00fea4d3e3b4Builds on5
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 51 citations
- Out-of-Distribution Generalization in Natural Language Processing: Past, Present, and FutureLinyi Yang, Yaoxian Song, Xuan Ren, Chenyang Lyu et al.EMNLP 2023 · 12 citations
- iSarcasm: A Dataset of Intended SarcasmSilviu Oprea, Walid MagdyACL 2020 · 2 citations
Related papers
- Can Third Parties Read Our Emotions?Jiayi Li, Yingfan Zhou, Pranav Narayanan Venkit, Halima Binte Islam et al.ACL 2025 · 5 citations
- Contextual Gaps in Machine Learning for Mental Illness Prediction: The Case of Diagnostic DisclosuresStevie Chancellor, Jessica L. Feuston, Jayhyun ChangCSCW 2023 · 6 citations
- Detecting Users' Emotional States during Passive Social Media UseChristoph Gebhardt, Andreas Brombach, Tiffany Luong, Otmar Hilliges et al.UbiComp 2024 · 8 citations
- "Finsta gets all my bad pictures": Instagram Users' Self-Presentation Across Finsta and Rinsta AccountsXiaoyun Huang, Jessica VitakCSCW 2022 · 39 citations
- Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text PerceptionsMatthias Orlikowski, Jiaxin Pei, Paul Röttger, Philipp Cimiano et al.ACL 2025 · 34 citations
