Dialect Diversity in Text Summarization on Twitter
Vijay Keswani, L. Elisa Celis
摘要
Discussions on Twitter involve participation from different communities with different dialects and it is often necessary to summarize a large number of posts into a representative sample to provide a synopsis. Yet, any such representative sample should sufficiently portray the underlying dialect diversity to present the voices of different participating communities representing the dialects. Extractive summarization algorithms perform the task of constructing subsets that succinctly capture the topic of any given set of posts. However, we observe that there is dialect bias in the summaries generated by common summarization approaches, i.e., they often return summaries that under-represent certain dialects. The vast majority of existing “fair” summarization approaches require socially salient attribute labels (in this case, dialect) to ensure that the generated summary is fair with respect to the socially salient attribute. Nevertheless, in many applications, these labels do not exist. Furthermore, due to the ever-evolving nature of dialects in social media, it is unreasonable to label or accurately infer the dialect of every social media post. To correct for the dialect bias, we employ a framework that takes an existing text summarization algorithm as a blackbox and, using a small set of dialect-diverse sentences, returns a summary that is relatively more dialect-diverse. Crucially, this approach does not need the posts being summarized to have dialect labels, ensuring that the diversification process is independent of dialect classification/identification models. We show the efficacy of our approach on Twitter datasets containing posts written in dialects used by different social groups defined by race or gender; in all cases, our approach leads to improved dialect diversity compared to standard text summarization approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Maximizing Submodular Functions for Recommendation in the Presence of BiasesAnay Mehrotra, Nisheeth K. VishnoiWWW 2023 · 被引用 11 次
- Summarizing Speech: A Comprehensive SurveyFabian Retkowski, Maike Züfle, Andreas Sudmann, Dinah Pfau 等EMNLP 2025 · 被引用 3 次
- Data Caricatures: On the Representation of African American Language in Pretraining CorporaNicholas Deas, Blake Vente, Amith Ananthram, Jessica Grieser 等ACL 2025
它引用的顶会 Paper5
- Extractive Summarization as Text MatchingMing Zhong, Pengfei Liu, Yiran Chen, Danqing Wang 等ACL 2020 · 被引用 410 次
- Fair Generative Modeling via Weak SupervisionKristy Choi, Aditya Grover, Trisha Singh, Rui Shu 等ICML 2020 · 被引用 160 次
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 被引用 68 次
- Implicit Diversity in Image SummarizationL. Elisa Celis, Vijay KeswaniCSCW 2020 · 被引用 26 次
- StereoSet: Measuring stereotypical bias in pretrained language modelsMoin Nadeem, Anna Bethke, Siva ReddyACL 2021
相关 Paper
- Auditing for Diversity Using Representative ExamplesVijay Keswani, L. Elisa CelisKDD 2021 · 被引用 1 次
- Neural Label Search for Zero-Shot Multi-Lingual Extractive SummarizationRuipeng Jia, Xingxing Zhang, Yanan Cao, Zheng Lin 等ACL 2022
- FuzzE: Fuzzy Fairness Evaluation of Offensive Language Classifiers on African-American EnglishAnthony RiosAAAI 2020 · 被引用 26 次
- Jointly Learning to Align and Summarize for Neural Cross-Lingual SummarizationYue Cao, Hui Liu, Xiaojun WanACL 2020 · 被引用 52 次
- Alt-Text with Context: Improving Accessibility for Images on TwitterNikita Srivatsan, Sofía Samaniego, Omar Florez, Taylor Berg-KirkpatrickICLR 2024 · 被引用 9 次
