NAT: Noise-Aware Training for Robust Neural Sequence Labeling
Marcin Namysl, Sven Behnke, Joachim Köhler
摘要
Sequence labeling systems should perform reliably not only under ideal conditions but also with corrupted inputs-as these systems often process user-generated text or follow an errorprone upstream component. To this end, we formulate the noisy sequence labeling problem, where the input may undergo an unknown noising process and propose two Noise-Aware Training (NAT) objectives that improve robustness of sequence labeling performed on perturbed input: Our data augmentation method trains a neural model using a mixture of clean and noisy samples, whereas our stability training algorithm encourages the model to create a noise-invariant latent representation. We employ a vanilla noise model at training time. For evaluation, we use both the original data and its variants perturbed with real OCR errors and misspellings. Extensive experiments on English and German named entity recognition benchmarks confirmed that NAT consistently improved robustness of popular sequence labeling models, preserving accuracy on the original input. We make our code and data publicly available for the research community.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- READIN: A Chinese Multi-Task Benchmark with Realistic and Diverse Input NoisesChenglei Si, Zhengyan Zhang, Yingfa Chen, Xiaozhi Wang 等ACL 2023 · 被引用 1 次
- Priority on High-Quality: Selecting Instruction Data via Consistency Verification of Noise InjectionHong Zhang, Feng Zhao, Ruilin Zhao, Cheng Yan 等EMNLP 2025
相关 Paper
- NoiseBench: Benchmarking the Impact of Real Label Noise on Named Entity RecognitionElena Merdjanovska, Ansar Aynetdinov, Alan AkbikEMNLP 2024 · 被引用 5 次
- Distantly-Supervised Named Entity Recognition with Noise-Robust Learning and Language Model Augmented Self-TrainingYu Meng, Yunyi Zhang, Jiaxin Huang, Xuan Wang 等EMNLP 2021 · 被引用 50 次
- Robust and Informative Text Augmentation (RITA) via Constrained Worst-Case Transformations for Low-Resource Named Entity RecognitionHyunwoo Sohn, Baekkwan ParkKDD 2022 · 被引用 3 次
- Uncertainty-Aware Self-Training for Low-Resource Neural Sequence LabelingJianing Wang, Chengyu Wang, Jun Huang, Ming Gao 等AAAI 2023 · 被引用 5 次
- Robust Named Entity Recognition with Truecasing PretrainingStephen Mayhew, Nitish Gupta, Dan RothAAAI 2020 · 被引用 43 次
