CS2W: A Chinese Spoken-to-Written Style Conversion Dataset with Multiple Conversion Types
Zishan Guo, Linhao Yu, Minghui Xu, Renren Jin, Deyi Xiong
Abstract
Spoken texts (either manual or automatic transcriptions from automatic speech recognition (ASR)) often contain disfluencies and grammatical errors, which pose tremendous challenges to downstream tasks. Converting spoken into written language is hence desirable. Unfortunately, the availability of datasets for this is limited. To address this issue, we present CS2W, a Chinese Spoken-to-Written style conversion dataset comprising 7,237 spoken sentences extracted from transcribed conversational texts. Four types of conversion problems are covered in CS2W: disfluencies, grammatical errors, ASR transcription errors, and colloquial words. Our annotation convention, data, and code are publicly available at https://github.com/guozishan/CS2W .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Recording for Eyes, Not Echoing to Ears: Contextualized Spoken-to-Written Conversion of ASR TranscriptsJiaqing Liu, Chong Deng, Qinglin Zhang, Shilin Zhou et al.AAAI 2025 · 1 citation
- GeWu: A Culturally-Grounded Chinese Benchmark for Multi-Stage Social Bias Evaluation in Large Language ModelsYi Lin, Ziyi Zhou, Jiashi Gao, Xinwei Guo et al.AAAI 2026
- COAS2W: A Chinese Older-Adults Spoken-to-Written Transformation Corpus with Context AwarenessChun Kang, Zhigu Qian, Zhen Fu, Jiaojiao Fu et al.EMNLP 2025
Builds on3
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Tail-to-Tail Non-Autoregressive Sequence Prediction for Chinese Grammatical Error CorrectionPiji Li, Shuming ShiACL 2021
- GLM: General Language Model Pretraining with Autoregressive Blank InfillingZhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding et al.ACL 2022
Related papers
- RealTalk-CN: A Realistic Chinese Speech Task-Oriented Dialogue Benchmark with Cross-Modal AnalysisEnzhi Wang, Jiaming Zhou, Yuhang Jia, Aobo Kong et al.ACL 2026
- UMRSpell: Unifying the Detection and Correction Parts of Pre-trained Models towards Chinese Missing, Redundant, and Spelling CorrectionZheyu He, Yujin Zhu, Linlin Wang, Liang XuACL 2023 · 8 citations
- Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and SentencesDmitrii Korzh, Dmitrii Tarasov, Artyom Iudin, Elvir Karimov et al.ICLR 2026 · 2 citations
- Why Aren't We NER Yet? Artifacts of ASR Errors in Named Entity Recognition in Spontaneous Speech TranscriptsPiotr Szymanski, Lukasz Augustyniak, Mikolaj Morzy, Adrian Szymczak et al.ACL 2023 · 7 citations
- Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-trainingYifan Yang, Bing Han, Hui Wang, Wei Wang et al.ACL 2026 · 4 citations
