Align With Purpose: Optimize Desired Properties in CTC Models with a General Plug-and-Play Framework
Eliya Segev, Maya Alroy, Ronen Katsir, Noam Wies, Ayana Shenhav, Yael Ben-Oren, David Zar, Oren Tadmor, Jacob Bitterman, Amnon Shashua, Tal Rosenwein
Abstract
Connectionist Temporal Classification (CTC) is a widely used criterion for training supervised sequence-to-sequence (seq2seq) models. It learns the alignments between the input and output sequences by marginalizing over the perfect alignments (that yield the ground truth), at the expense of the imperfect ones. This dichotomy, and in particular the equal treatment of all perfect alignments, results in a lack of controllability over the predicted alignments. This controllability is essential for capturing properties that hold significance in real-world applications. Here we propose Align With Purpose (AWP), a general Plug-and-Play framework for enhancing a desired property in models trained with the CTC criterion. We do that by complementing the CTC loss with an additional loss term that prioritizes alignments according to a desired property. AWP does not require any intervention in the CTC loss function, and allows to differentiate between both perfect and imperfect alignments for a variety of properties. We apply our framework in the domain of Automatic Speech Recognition (ASR) and show its generality in terms of property selection, architectural choice, and scale of the training dataset (up to 280,000 hours). To demonstrate the effectiveness of our framework, we apply it to two unrelated properties: token emission time for latency optimization and word error rate (WER). For the former, we report an improvement of up to 590ms in latency optimization with a minor reduction in WER, and for the latter, we report a relative improvement of 4.5% in WER over the baseline models. To the best of our knowledge, these applications have never been demonstrated to work on this scale of data. Notably, our method can be easily implemented using only a few lines of code 1 and can be extended to other alignment-free loss functions and to domains other than ASR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fee324bc-d4ee-48b2-91fd-3b861ee0903bCited by top-tier papers1
Ask how each one uses itBuilds on6
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen et al.ICLR 2020 · 2,210 citations
- Non-autoregressive Translation with Layer-Wise Prediction and Deep SupervisionChenyang Huang, Hao Zhou, Osmar R. Zaïane, Lili Mou et al.AAAI 2022 · 65 citations
- Bayes Risk CTC: Controllable CTC Alignment in Sequence-to-Sequence TasksJinchuan Tian, Brian Yan, Jianwei Yu, Chao Weng et al.ICLR 2023 · 4 citations
Related papers
- W-CTC: a Connectionist Temporal Classification Loss with Wild CardsXingyu Cai, Jiahong Yuan, Yuchen Bian, Guangxu Xun et al.ICLR 2022 · 16 citations
- CR-CTC: Consistency regularization on CTC for improved speech recognitionZengwei Yao, Wei Kang, Xiaoyu Yang, Fangjun Kuang et al.ICLR 2025
- Uncertainty-Aware Self-Training for CTC-Based Automatic Speech RecognitionEungbeom Kim, Kyogu LeeAAAI 2025 · 2 citations
- SoftCorrect: Error Correction with Soft Detection for Automatic Speech RecognitionYichong Leng, Xu Tan, Wenjie Liu, Kaitao Song et al.AAAI 2023 · 22 citations
- Star Temporal Classification: Sequence Modeling with Partially Labeled DataVineel Pratap, Awni Hannun, Gabriel Synnaeve, Ronan CollobertNeurIPS 2022 · 7 citations
