Towards Author-informed NLP: Mind the Social Bias
Inbar Pendzel, Einat Minkov
Abstract
Social text understanding is prone to fail when opinions are conveyed implicitly or sarcastically. It is therefore desired to model users' contexts in processing the texts authored by them. In this work, we represent users within a social embedding space that was learned from the Twitter network at large-scale. Similar to word embeddings that encode lexical semantics, the network embeddings encode latent dimensions of social semantics. We perform extensive experiments on author-informed stance prediction, demonstrating improved generalization through inductive social user modeling, both within and across topics. Similar results were obtained for author-informed toxicity and incivility detection. The proposed approach may pave way to social NLP that considers user embeddings as contextual modality. However, our investigation also reveals that user stances are correlated with the personal socio-demographic traits encoded in their embeddings. Hence, author-informed NLP approaches may inadvertently model and reinforce socio-demographic and other social biases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 392d0ef6-28fd-4ae2-8be8-e4b3e5b46495Builds on3
- Demographic Representation and Collective Storytelling in the Me Too Twitter Hashtag Activism MovementAaron Mueller, Zach Wood-Doughty, Silvio Amir, Mark Dredze et al.CSCW 2021 · 53 citations
- A Closer Look at Multidimensional Online Political IncivilitySagi Pendzel, Nir Lotan, Alon Zoizner, Einat MinkovEMNLP 2024 · 2 citations
- Multimodal Transformers are Hierarchical Modal-wise Heterogeneous GraphsYijie Jin, Junjie Peng, Xuanchao Lin, Haochen Yuan et al.ACL 2025
Related papers
- An Embedding Model for Estimating Legislative Preferences from the Frequency and Sentiment of TweetsGregory Spell, Brian Guay, Sunshine Hillygus, Lawrence CarinEMNLP 2020 · 6 citations
- The Structure of Toxic Conversations on TwitterMartin Saveski, Brandon Roy, Deb RoyWWW 2021 · 111 citations
- Unsupervised Detection of Contextualized Embedding Bias with Application to IdeologyValentin Hofmann, Janet B. Pierrehumbert, Hinrich SchützeICML 2022 · 1 citation
- Automated Detection of Doxing on TwitterYounes Karimi, Anna Cinzia Squicciarini, Shomir WilsonCSCW 2022 · 18 citations
- Are Stereotypes Leading LLMs' Zero-Shot Stance Detection ?Anthony Dubreuil, Antoine Gourru, Christine Largeron, Amine TrabelsiEMNLP 2025 · 1 citation
