Digital Skin, Digital Bias: Uncovering Tone-Based Biases in LLMs and Emoji Embeddings
Mingchen Li, Wajdi Aljedaani, Yingjie Liu, Navyasri Meka, Xuan Lu, Xinyue Ye, Junhua Ding, Yunhe Feng
Abstract
Skin-toned emojis are crucial for fostering personal identity and social inclusion in online communication. As AI models, particularly Large Language Models (LLMs), increasingly mediate interactions on web platforms, the risk that these systems perpetuate societal biases through their representation of such symbols is a significant concern. This paper presents the first large-scale comparative study of bias in skin-toned emoji representations across two distinct model classes. We systematically evaluate dedicated emoji embedding models (emoji2vec, emoji-sw2v) against four modern LLMs (Llama, Gemma, Qwen, and Mistral). Our analysis first reveals a critical performance gap: while LLMs demonstrate robust support for skin tone modifiers, widely-used specialized emoji models exhibit severe deficiencies. More importantly, a multi-faceted investigation into semantic consistency, representational similarity, sentiment polarity, and core biases uncovers systemic disparities. We find evidence of skewed sentiment and inconsistent meanings associated with emojis across different skin tones, highlighting latent biases within these foundational models. Our findings underscore the urgent need for developers and platforms to audit and mitigate these representational harms, ensuring that AI's role on the web promotes genuine equity rather than reinforcing societal biases. CCS Concepts • Computing methodologies → Natural language processing; Machine learning; • Information systems → World Wide Web; • Social and professional topics → Race and ethnicity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ea394d72-2866-4e99-86f8-54600ae4a6eaRelated papers
- EMODIS: A Benchmark for Context-Dependent Emoji Disambiguation in Large Language ModelsJiacheng Huang, Ning Yu, Xiaoyin YiAAAI 2026
- Job Unfair: An Investigation of Gender and Occupational Bias in Free-Form Text Completions by LLMsCamilla Casula, Sebastiano Vecellio Salto, Elisa Leonardelli, Sara TonelliEMNLP 2025
- VisBias: Measuring Explicit and Implicit Social Biases in Vision Language ModelsJen-Tse Huang, Jiantong Qin, Jianping Zhang, Youliang Yuan et al.EMNLP 2025 · 13 citations
- Bias Similarity Measurement: A Black-Box Audit of Fairness Across LLMsHyejun Jeong, Shiqing Ma, Amir HoumansadrICLR 2026 · 1 citation
- Angry Men, Sad Women: Large Language Models Reflect Gendered Stereotypes in Emotion AttributionFlor Miriam Plaza del Arco, Amanda Cercas Curry, Alba Cercas Curry, Gavin Abercrombie et al.ACL 2024 · 8 citations
