LLM-ForcedAligner: A Non-Autoregressive and Accurate LLM-Based Forced Aligner for Multilingual and Long-Form Speech
Bingshen Mu, Xian Shi, Xiong Wang, Hexin Liu, Jin Xu, Lei Xie
Abstract
Forced alignment (FA) predicts start and end timestamps for words or characters in speech, but existing methods are language-specific and prone to cumulative temporal shifts. The multilingual speech understanding and long-sequence processing abilities of speech large language models (SLLMs) make them promising for FA in multilingual, crosslingual, and long-form speech settings. However, directly applying the next-token prediction paradigm of SLLMs to FA results in hallucinations and slow inference. To bridge the gap, we propose LLM-ForcedAligner, reformulating FA as a slot-filling paradigm: timestamps are treated as discrete indices, and special timestamp tokens are inserted as slots into the transcript. Conditioned on the speech embeddings and the transcript with slots, the SLLM directly predicts the time indices at slots. During training, causal attention masking with non-shifted input and label sequences allows each slot to predict its own timestamp index based on itself and preceding context, with loss computed only at slot positions. Dynamic slot insertion enables FA at arbitrary positions. Moreover, non-autoregressive inference is supported, avoiding hallucinations and improving speed. Experiments across multilingual, crosslingual, and long-form speech scenarios show that LLM-ForcedAligner achieves a 69% 78% relative reduction in accumulated averaging shift compared with prior methods. Checkpoint and inference code are available at https://github.com/QwenLM/Qwen3-ASR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0c0a11bc-4c28-4ea9-a26b-0e7b893682e3Builds on1
Related papers
- AutoTimes: Autoregressive Time Series Forecasters via Large Language ModelsYong Liu, Guo Qin, Xiangdong Huang, Jianmin Wang et al.NeurIPS 2024 · 138 citations
- TimeCMA: Towards LLM-Empowered Multivariate Time Series Forecasting via Cross-Modality AlignmentChenxi Liu, Qianxiong Xu, Hao Miao, Sun Yang et al.AAAI 2025 · 141 citations
- Markovian Linguistic-Temporal Bridge: Unlocking the Potential of LLMs for Time Series ForecastingSiming Sun, Kai Zhang, Xuejun Jiang, Wenchao Meng et al.ACL 2026
- Introducing Semantics into Speech EncodersDerek Xu, Shuyan Dong, Changhan Wang, Suyoun Kim et al.ACL 2023 · 2 citations
- AlignCap: Aligning Speech Emotion Captioning to Human PreferencesZiqi Liang, Haoxiang Shi, Hanhui ChenEMNLP 2024 · 2 citations
