Dialect Diversity in Text Summarization on Twitter
Vijay Keswani, L. Elisa Celis
Abstract
Discussions on Twitter involve participation from different communities with different dialects and it is often necessary to summarize a large number of posts into a representative sample to provide a synopsis. Yet, any such representative sample should sufficiently portray the underlying dialect diversity to present the voices of different participating communities representing the dialects. Extractive summarization algorithms perform the task of constructing subsets that succinctly capture the topic of any given set of posts. However, we observe that there is dialect bias in the summaries generated by common summarization approaches, i.e., they often return summaries that under-represent certain dialects. The vast majority of existing “fair” summarization approaches require socially salient attribute labels (in this case, dialect) to ensure that the generated summary is fair with respect to the socially salient attribute. Nevertheless, in many applications, these labels do not exist. Furthermore, due to the ever-evolving nature of dialects in social media, it is unreasonable to label or accurately infer the dialect of every social media post. To correct for the dialect bias, we employ a framework that takes an existing text summarization algorithm as a blackbox and, using a small set of dialect-diverse sentences, returns a summary that is relatively more dialect-diverse. Crucially, this approach does not need the posts being summarized to have dialect labels, ensuring that the diversification process is independent of dialect classification/identification models. We show the efficacy of our approach on Twitter datasets containing posts written in dialects used by different social groups defined by race or gender; in all cases, our approach leads to improved dialect diversity compared to standard text summarization approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e3fc85a3-e268-449a-bbcf-2ee3a9f66242Cited by top-tier papers3
- Maximizing Submodular Functions for Recommendation in the Presence of BiasesAnay Mehrotra, Nisheeth K. VishnoiWWW 2023 · 11 citations
- Summarizing Speech: A Comprehensive SurveyFabian Retkowski, Maike Züfle, Andreas Sudmann, Dinah Pfau et al.EMNLP 2025 · 3 citations
- Data Caricatures: On the Representation of African American Language in Pretraining CorporaNicholas Deas, Blake Vente, Amith Ananthram, Jessica Grieser et al.ACL 2025
Builds on5
- Extractive Summarization as Text MatchingMing Zhong, Pengfei Liu, Yiran Chen, Danqing Wang et al.ACL 2020 · 410 citations
- Fair Generative Modeling via Weak SupervisionKristy Choi, Aditya Grover, Trisha Singh, Rui Shu et al.ICML 2020 · 160 citations
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 68 citations
- Implicit Diversity in Image SummarizationL. Elisa Celis, Vijay KeswaniCSCW 2020 · 26 citations
- StereoSet: Measuring stereotypical bias in pretrained language modelsMoin Nadeem, Anna Bethke, Siva ReddyACL 2021
Related papers
- Auditing for Diversity Using Representative ExamplesVijay Keswani, L. Elisa CelisKDD 2021 · 1 citation
- Neural Label Search for Zero-Shot Multi-Lingual Extractive SummarizationRuipeng Jia, Xingxing Zhang, Yanan Cao, Zheng Lin et al.ACL 2022
- FuzzE: Fuzzy Fairness Evaluation of Offensive Language Classifiers on African-American EnglishAnthony RiosAAAI 2020 · 26 citations
- Jointly Learning to Align and Summarize for Neural Cross-Lingual SummarizationYue Cao, Hui Liu, Xiaojun WanACL 2020 · 52 citations
- Alt-Text with Context: Improving Accessibility for Images on TwitterNikita Srivatsan, Sofía Samaniego, Omar Florez, Taylor Berg-KirkpatrickICLR 2024 · 9 citations
