Sequence to Multi-Sequence Learning via Conditional Chain Mapping for Mixture Signals
Jing Shi, Xuankai Chang, Pengcheng Guo, Shinji Watanabe, Yusuke Fujita, Jiaming Xu, Bo Xu, Lei Xie
Abstract
Neural sequence-to-sequence models are well established for applications which can be cast as mapping a single input sequence into a single output sequence. In this work, we focus on one-to-many sequence transduction problems, such as extracting multiple sequential sources from a mixture sequence. We extend the standard sequence-to-sequence model to a conditional multi-sequence model, which explicitly models the relevance between multiple output sequences with the probabilistic chain rule. Based on this extension, our model can conditionally infer output sequences one-by-one by making use of both input and previously-estimated contextual output sequences. This model additionally has a simple and efficient stop criterion for the end of the transduction, making it able to infer the variable number of output sequences. We take speech data as a primary test field to evaluate our methods since the observed speech data is often composed of multiple sources due to the nature of the superposition principle of sound waves. Experiments on several different tasks including speech separation and multi-speaker speech recognition show that our conditional multi-sequence models lead to consistent improvements over the conventional non-conditional models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c1aa55cc-3597-4a19-8269-ff5e526bb841Builds on1
Related papers
- LCMA-SRT: Language-Conditional Mixture-of-Experts Adapters for Joint Multilingual Speech Recognition and TranslationNanjie Li, Xiaoyong Guo, Hao Huang, Haihua Xu et al.ACL 2026
- Copy That! Editing Sequences by Copying SpansSheena Panthaplackel, Miltiadis Allamanis, Marc BrockschmidtAAAI 2021 · 28 citations
- Mu2SLAM: Multitask, Multilingual Speech and Language ModelsYong Cheng, Yu Zhang, Melvin Johnson, Wolfgang Macherey et al.ICML 2023 · 10 citations
- MixSeq: Connecting Macroscopic Time Series Forecasting with Microscopic Time Series DataZhibo Zhu, Ziqi Liu, Ge Jin, Zhiqiang Zhang et al.NeurIPS 2021 · 16 citations
- Audio2Gestures: Generating Diverse Gestures from Speech Audio with Conditional Variational AutoencodersJing Li, Di Kang, Wenjie Pei, Xuefei Zhe et al.ICCV 2021 · 144 citations
