Code-switched inspired losses for spoken dialog representations
Pierre Colombo, Emile Chapuis, Matthieu Labeau, Chloé Clavel
摘要
Spoken dialog systems need to be able to handle both multiple languages and multilinguality inside a conversation (e.g in case of codeswitching). In this work, we introduce new pretraining losses tailored to learn multilingual spoken dialog representations. The goal of these losses is to expose the model to codeswitched language. To scale up training, we automatically build a pretraining corpus composed of multilingual conversations in five different languages (French, Italian, English, German and Spanish) from OpenSubtitles, a huge multilingual corpus composed of 24.3G tokens. We test the generic representations on MIAM, a new benchmark composed of five dialog act corpora on the same aforementioned languages as well as on two novel multilingual downstream tasks (i.e multilingual mask utterance retrieval and multilingual inconsistency identification). Our experiments show that our new code switched-inspired losses achieve a better performance in both monolingual and multilingual settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Beyond Mahalanobis Distance for Textual OOD DetectionPierre Colombo, Eduardo Dadalto Câmara Gomes, Guillaume Staerman, Nathan Noiry 等NeurIPS 2022 · 被引用 24 次
- Transductive Learning for Textual Few-Shot Classification in API-based Embedding ModelsPierre Colombo, Victor Pellegrain, Malik Boudiaf, Myriam Tami 等EMNLP 2023 · 被引用 7 次
- Towards Zero-Shot Multilingual Transfer for Code-Switched ResponsesTing-Wei Wu, Changsheng Zhao, Ernie Chang, Yangyang Shi 等ACL 2023 · 被引用 2 次
- Learning Disentangled Textual Representations via Statistical Measures of SimilarityPierre Colombo, Guillaume Staerman, Nathan Noiry, Pablo PiantanidaACL 2022
它引用的顶会 Paper10
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Cross-Lingual Ability of Multilingual BERT: An Empirical StudyKarthikeyan K, Zihan Wang, Stephen Mayhew, Dan RothICLR 2020 · 被引用 378 次
- Guiding Attention in Sequence-to-Sequence Models for Dialogue Act PredictionPierre Colombo, Emile Chapuis, Matteo Manica, Emmanuel Vignon 等AAAI 2020 · 被引用 69 次
相关 Paper
- M3P: Learning Universal Representations via Multitask Multilingual Multimodal Pre-TrainingMinheng Ni, Haoyang Huang, Lin Su, Edward Cui 等CVPR 2021
- The Role of Mixed-Language Documents for Multilingual Large Language Model PretrainingJiandong Shao, Raphael Tang, Crystina Zhang, Karin Sevegnani 等ACL 2026
- Alternating Language Modeling for Cross-Lingual Pre-TrainingJian Yang, Shuming Ma, Dongdong Zhang, Shuangzhi Wu 等AAAI 2020 · 被引用 94 次
- Attention-Informed Mixed-Language Training for Zero-Shot Cross-Lingual Task-Oriented Dialogue SystemsZihan Liu, Genta Indra Winata, Zhaojiang Lin, Peng Xu 等AAAI 2020 · 被引用 105 次
- Cross-lingual Intermediate Fine-tuning improves Dialogue State TrackingNikita Moghe, Mark Steedman, Alexandra BirchEMNLP 2021 · 被引用 11 次
