GupShup: Summarizing Open-Domain Code-Switched Conversations
Laiba Mehnaz, Debanjan Mahata, Rakesh Gosangi, Uma Sushmitha Gunturi, Riya Jain, Gauri Gupta, Amardeep Kumar, Isabelle G. Lee, Anish Acharya, Rajiv Ratn Shah
摘要
Code-switching is the communication phenomenon where the speakers switch between different languages during a conversation. With the widespread adoption of conversational agents and chat platforms, codeswitching has become an integral part of written conversations in many multi-lingual communities worldwide. Therefore, it is essential to develop techniques for understanding and summarizing these conversations. Towards this objective, we introduce the task of abstractive summarization of Hindi-English (Hi-En) code-switched conversations. We also develop the first code-switched conversation summarization dataset -GupShup, which contains over 6,800 Hi-En conversations and their corresponding human-annotated summaries in English (En) and Hi-En. We present a detailed account of the entire data collection and annotation process. We analyze the dataset using various code-switching statistics. We train state-of-the-art abstractive summarization models and report their performances using both automated metrics and human evaluation. Our results show that multi-lingual mBART and multi-view seq2seq models obtain the best performances on this new dataset. We also conduct an extensive qualitative analysis to provide insight into the models and some of their shortcomings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ClidSum: A Benchmark Dataset for Cross-Lingual Dialogue SummarizationJiaan Wang, Fandong Meng, Ziyao Lu, Duo Zheng 等EMNLP 2022 · 被引用 27 次
- Multilingual Large Language Models Are Not (Yet) Code-SwitchersRuochen Zhang, Samuel Cahyawijaya, Jan Christian Blaise Cruz, Genta Indra Winata 等EMNLP 2023 · 被引用 19 次
- Beyond Monolingual Assumptions: A Survey on Code-Switched NLP in the Era of Large Language Models across ModalitiesRajvee Sheth, Samridhi Raj Sinha, Mahavir Patil, Himanshu Beniwal 等ACL 2026 · 被引用 2 次
它引用的顶会 Paper4
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Multi-View Sequence-to-Sequence Models with Conversational Structure for Abstractive Dialogue SummarizationJiaao Chen, Diyi YangEMNLP 2020 · 被引用 121 次
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 被引用 40 次
- The State and Fate of Linguistic Diversity and Inclusion in the NLP WorldPratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali 等ACL 2020 · 被引用 40 次
相关 Paper
- From Machine Translation to Code-Switching: Generating High-Quality Code-Switched TextIshan Tarunesh, Syamantak Kumar, Preethi JyothiACL 2021
- GLUECoS: An Evaluation Benchmark for Code-Switched NLPSimran Khanuja, Sandipan Dandapat, Anirudh Srinivasan, Sunayana Sitaram 等ACL 2020 · 被引用 12 次
- CoCoa: An Encoder-Decoder Model for Controllable Code-switched GenerationSneha Mondal, Ritika, Shreya Pathak, Preethi Jyothi 等EMNLP 2022 · 被引用 5 次
- From English to Code-Switching: Transfer Learning with Strong Morphological CluesGustavo Aguilar, Thamar SolorioACL 2020 · 被引用 1 次
- An Empirical Study of Many-to-Many Summarization with Large Language ModelsJiaan Wang, Fandong Meng, Zengkui Sun, Yunlong Liang 等ACL 2025
