Perception and Semantic Aware Regularization for Sequential Confidence Calibration
Zhenghua Peng, Yu Luo, Tianshui Chen, Keke Xu, Shuangping Huang
Abstract
Deep sequence recognition (DSR) models receive increasing attention due to their superior application to various applications. Most DSR models use merely the target sequences as supervision without considering other related sequences, leading to over-confidence in their predictions. The DSR models trained with label smoothing regularize labels by equally and independently smoothing each token, reallocating a small value to other tokens for mitigating overconfidence. However, they do not consider tokens/sequences correlations that may provide more effective information to regularize training and thus lead to sub-optimal performance. In this work, we find tokens/sequences with high perception and semantic correlations with the target ones contain more correlated and effective information and thus facilitate more effective regularization. To this end, we propose a Perception and Semantic aware Sequence Regularization framework, which explore perceptively and semantically correlated tokens/sequences as regularization. Specifically, we introduce a semantic context-free recognition and a language model to acquire similar sequences with high perceptive similarities and semantic correlation, respectively. Moreover, over-confidence degree varies across samples according to their difficulties. Thus, we further design an adaptive calibration intensity module to compute a difficulty score for each samples to obtain finer-grained regularization. Extensive experiments on canonical sequence recognition tasks, including scene text and speech recognition, demonstrate that our method sets novel state-of-the-art results. Code is available at https://github.com/husterpzh/PSSR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 313eba82-393b-4512-aa0e-58000602c5a6Builds on35
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- Confidence Regularized Self-TrainingYang Zou, Zhiding Yu, Xiaofeng Liu, B. V. K. Vijaya Kumar et al.ICCV 2019 · 901 citations
- Conformer: Local Features Coupling Global Representations for Visual RecognitionZhiliang Peng, Wei Huang, Shanzhi Gu, Lingxi Xie et al.ICCV 2021 · 723 citations
- Calibrating Deep Neural Networks using Focal LossJishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz et al.NeurIPS 2020 · 674 citations
Related papers
- Context-Aware Selective Label Smoothing for Calibrating Sequence Recognition ModelShuangping Huang, Yu Luo, Zhenzhou Zhuang, Jin-Gang Yu et al.ACM MM 2021 · 10 citations
- Trust Prophet or Not? Taking a Further Verification Step toward Accurate Scene Text RecognitionAnna Zhu, Ke Xiao, Bo Zhou, Runmin WangACM MM 2024 · 6 citations
- Pushing the Performance Limit of Scene Text Recognizer without Human AnnotationCaiyuan Zheng, Hui Li, Seon-Min Rhee, Seungju Han et al.CVPR 2022 · 20 citations
- Fine-grained Pseudo Labels for Scene Text RecognitionXiaoyu Li, Xiaoxue Chen, Zuming Huang, Lele Xie et al.ACM MM 2023 · 2 citations
- SCATTER: Selective Context Attentional Scene Text RecognizerRon Litman, Oron Anschel, Shahar Tsiper, Roee Litman et al.CVPR 2020
