ViDia2Std: A Parallel Corpus and Methods for Low-Resource Vietnamese Dialect-to-Standard Translation
Khoa Anh Ta, Nguyen Van Dinh, Kiet Van Nguyen
Abstract
Vietnamese exhibits extensive dialectal variation, posing challenges for NLP systems trained predominantly on standard Vietnamese. Such systems often underperform on dialectal inputs, especially from underrepresented Central and Southern regions. Previous work on dialect normalization has focused narrowly on Central-to-Northern dialect transfer using synthetic data and limited dialectal diversity. These efforts exclude Southern varieties and intra-regional variants within the North. We introduce ViDia2Std, the first manually annotated parallel corpus for dialect-to-standard Vietnamese translation covering all 63 provinces. Unlike prior datasets, ViDia2Std includes diverse dialects from Central, Southern, and non-standard Northern regions often absent from existing resources, making it the most dialectally inclusive corpus to date. The dataset consists of over 13,000 sentence pairs sourced from real-world Facebook comments and annotated by native speakers across all three dialect regions. To assess annotation consistency, we define a semantic mapping agreement metric that accounts for synonymous standard mappings across annotators. Based on this criterion, we report agreement rates of 86% (North), 82% (Central), and 85% (South). We benchmark several sequence-to-sequence models on ViDia2Std. mBART-large-50 achieves the best results (BLEU 0.8166, ROUGE-L 0.9384, METEOR 0.8925), while ViT5-base offers competitive performance with fewer parameters. ViDia2Std demonstrates that dialect normalization substantially improves downstream tasks, highlighting the need for dialect-aware resources in building robust Vietnamese NLP systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a0407fe0-df54-4a40-81a3-eadd38ca4488Related papers
- Multi-Dialect Vietnamese: Task, Dataset, Baseline Models and ChallengesNguyen Dinh, Thanh Dang, Luan Thanh Nguyen, Kiet Van NguyenEMNLP 2024 · 2 citations
- DIALECTBENCH: An NLP Benchmark for Dialects, Varieties, and Closely-Related LanguagesFahim Faisal, Orevaoghene Ahia, Aarohi Srivastava, Kabir Ahuja et al.ACL 2024 · 10 citations
- MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech TranslationKhai Le-Duc, Tuyen Tran, Bach Phan Tat, Nguyen Kim Hai Bui et al.EMNLP 2025
- ARBERT & MARBERT: Deep Bidirectional Transformers for ArabicMuhammad Abdul-Mageed, AbdelRahim A. Elmadany, El Moatez Billah NagoudiACL 2021
- Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMsAbdellah El Mekki, Samar Mohamed Magdy, Houdaifa Atou, Ruwa AbuHweidi et al.ACL 2026
