Semi-supervised Contextual Historical Text Normalization
Peter Makarov, Simon Clematide
2020Year
7Citations
Abstract
. Yet, virtually all approaches suffer from the two limitations: 1) They consider a fully supervised setup, often with impractically large manually normalized datasets; 2) Normalization happens on words in isolation. By utilizing a simple generative normalization model and obtaining powerful contextualization from the target-side language model, we train accurate models with unlabeled historical data. In realistic training scenarios, our approach often leads to reduction in manually normalized data at the same accuracy levels.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Related papers
- Empirical Analysis of Unlabeled Entity Problem in Named Entity RecognitionYangming Li, Lemao Liu, Shuming ShiICLR 2021 · 72 citations
- You Are What You Train: Effects of Data Composition on Training Context-aware Machine Translation ModelsPawel Maka, Yusuf Can Semerci, Jan Scholtes, Gerasimos SpanakisEMNLP 2025
- Text Classification Using Label Names Only: A Language Model Self-Training ApproachYu Meng, Yunyi Zhang, Jiaxin Huang, Chenyan Xiong et al.EMNLP 2020 · 203 citations
- Handling Rare Entities for Neural Sequence LabelingYangming Li, Han Li, Kaisheng Yao, Xiaolong LiACL 2020 · 14 citations
- Normalization of Language Embeddings for Cross-Lingual AlignmentPrince Osei Aboagye, Yan Zheng, Chin-Chia Michael Yeh, Junpeng Wang et al.ICLR 2022 · 12 citations
