A Dataset and Baselines for Multilingual Reply Suggestion
Mozhi Zhang, Wei Wang, Budhaditya Deb, Guoqing Zheng, Milad Shokouhi, Ahmed Hassan Awadallah
Abstract
Reply suggestion models help users process emails and chats faster. Previous work only studies English reply suggestion. Instead, we present MRS, a multilingual reply suggestion dataset with ten languages. MRS can be used to compare two families of models: 1) retrieval models that select the reply from a fixed set and 2) generation models that produce the reply from scratch. Therefore, MRS complements existing cross-lingual generalization benchmarks that focus on classification and sequence labeling tasks. We build a generation model and a retrieval model as baselines for MRS. The two models have different strengths in the monolingual setting, and they require different strategies to generalize across languages. MRS is publicly available at https://github.com/zhangmozhi/mrs . Multilingual Reply Suggestion Automated reply suggestion is a useful feature for email and chat applications. Given an input message, the system suggests several replies, and users may click on them to save typing time (Figure 1 ). This feature is available in many applications including Gmail, Outlook, LinkedIn, Facebook Messenger, Microsoft Teams, and Uber. Reply suggestion is related to but different from open-domain dialog systems or chatbots (Adiwardana et al., 2020; Huang et al., 2020) . While both are conversational AI tasks (Gao et al., 2019) , the goals are different: reply suggestion systems help the user quickly reply to a message, while chatbots aim to continue the conversation and focus more on multi-turn dialogues. Ideally, we want our model to generate replies in any language. However, reply suggestion models require large training sets, so previous work mostly * Work mostly done as an intern at Microsoft Research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97890f33-ffc1-4e5e-b1af-0013a69764aaCited by top-tier papers1
Ask how each one uses itBuilds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan et al.ICLR 2020 · 683 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- XGLUE: A New Benchmark Datasetfor Cross-lingual Pre-training, Understanding and GenerationYaobo Liang, Nan Duan, Yeyun Gong, Ning Wu et al.EMNLP 2020 · 232 citations
Related papers
- XDailyDialog: A Multilingual Parallel Dialogue CorpusZeming Liu, Ping Nie, Jie Cai, Haifeng Wang et al.ACL 2023 · 4 citations
- M-RewardBench: Evaluating Reward Models in Multilingual SettingsSrishti Gureja, Lester James Validad Miranda, Shayekh Bin Islam, Rishabh Maheshwary et al.ACL 2025
- Understanding and Supporting Formal Email Exchange by Answering AI-Generated QuestionsYusuke Miura, Chi-Lan Yang, Masaki Kuribayashi, Keigo Matsumoto et al.CHI 2025 · 4 citations
- Comparing Sentence-Level Suggestions to Message-Level Suggestions in AI-Mediated CommunicationLiye Fu, Benjamin Newman, Maurice Jakesch, Sarah KrepsCHI 2023 · 28 citations
- MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media TextsDominik Macko, Jakub Kopal, Róbert Móro, Ivan SrbaACL 2025 · 15 citations
