Code-switched inspired losses for spoken dialog representations
Pierre Colombo, Emile Chapuis, Matthieu Labeau, Chloé Clavel
Abstract
Spoken dialog systems need to be able to handle both multiple languages and multilinguality inside a conversation (e.g in case of codeswitching). In this work, we introduce new pretraining losses tailored to learn multilingual spoken dialog representations. The goal of these losses is to expose the model to codeswitched language. To scale up training, we automatically build a pretraining corpus composed of multilingual conversations in five different languages (French, Italian, English, German and Spanish) from OpenSubtitles, a huge multilingual corpus composed of 24.3G tokens. We test the generic representations on MIAM, a new benchmark composed of five dialog act corpora on the same aforementioned languages as well as on two novel multilingual downstream tasks (i.e multilingual mask utterance retrieval and multilingual inconsistency identification). Our experiments show that our new code switched-inspired losses achieve a better performance in both monolingual and multilingual settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af05c7cd-60ff-459c-91fc-d63269f0a445Cited by top-tier papers4
- Beyond Mahalanobis Distance for Textual OOD DetectionPierre Colombo, Eduardo Dadalto Câmara Gomes, Guillaume Staerman, Nathan Noiry et al.NeurIPS 2022 · 24 citations
- Transductive Learning for Textual Few-Shot Classification in API-based Embedding ModelsPierre Colombo, Victor Pellegrain, Malik Boudiaf, Myriam Tami et al.EMNLP 2023 · 7 citations
- Towards Zero-Shot Multilingual Transfer for Code-Switched ResponsesTing-Wei Wu, Changsheng Zhao, Ernie Chang, Yangyang Shi et al.ACL 2023 · 2 citations
- Learning Disentangled Textual Representations via Statistical Measures of SimilarityPierre Colombo, Guillaume Staerman, Nathan Noiry, Pablo PiantanidaACL 2022
Builds on10
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Cross-Lingual Ability of Multilingual BERT: An Empirical StudyKarthikeyan K, Zihan Wang, Stephen Mayhew, Dan RothICLR 2020 · 378 citations
- Guiding Attention in Sequence-to-Sequence Models for Dialogue Act PredictionPierre Colombo, Emile Chapuis, Matteo Manica, Emmanuel Vignon et al.AAAI 2020 · 69 citations
Related papers
- M3P: Learning Universal Representations via Multitask Multilingual Multimodal Pre-TrainingMinheng Ni, Haoyang Huang, Lin Su, Edward Cui et al.CVPR 2021
- The Role of Mixed-Language Documents for Multilingual Large Language Model PretrainingJiandong Shao, Raphael Tang, Crystina Zhang, Karin Sevegnani et al.ACL 2026
- Alternating Language Modeling for Cross-Lingual Pre-TrainingJian Yang, Shuming Ma, Dongdong Zhang, Shuangzhi Wu et al.AAAI 2020 · 94 citations
- Attention-Informed Mixed-Language Training for Zero-Shot Cross-Lingual Task-Oriented Dialogue SystemsZihan Liu, Genta Indra Winata, Zhaojiang Lin, Peng Xu et al.AAAI 2020 · 105 citations
- Cross-lingual Intermediate Fine-tuning improves Dialogue State TrackingNikita Moghe, Mark Steedman, Alexandra BirchEMNLP 2021 · 11 citations
