Understanding Adaptive, Multiscale Temporal Integration In Deep Speech Recognition Systems
Menoua Keshishian, Samuel Norman-Haignere, Nima Mesgarani
Abstract
Natural signals such as speech are hierarchically structured across many different timescales, spanning tens (e.g., phonemes) to hundreds (e.g., words) of milliseconds, each of which is highly variable and context-dependent. While deep neural networks (DNNs) excel at recognizing complex patterns from natural signals, relatively little is known about how DNNs flexibly integrate across multiple timescales. Here, we show how a recently developed method for studying temporal integration in biological neural systems – the temporal context invariance (TCI) paradigm – can be used to understand temporal integration in DNNs. The method is simple: we measure responses to a large number of stimulus segments presented in two different contexts and estimate the smallest segment duration needed to achieve a context invariant response. We applied our method to understand how the popular DeepSpeech2 model learns to integrate across time in speech. We find that nearly all of the model units, even in recurrent layers, have a compact integration window within which stimuli substantially alter the response and outside of which stimuli have little effect. We show that training causes these integration windows to shrink at early layers and expand at higher layers, creating a hierarchy of integration windows across the network. Moreover, by measuring integration windows for time-stretched/compressed speech, we reveal a transition point, midway through the trained network, where integration windows become yoked to the duration of stimulus structures (e.g., phonemes or words) rather than absolute time. Similar phenomena were observed in a purely recurrent and purely convolutional network although structure-yoked integration was more prominent in the recurrent network. These findings suggest that deep speech recognition systems use a common motif to encode the hierarchical structure of speech: integrating across short, time-yoked windows at early layers and long, structure-yoked windows at later layers. Our method provides a straightforward and general-purpose toolkit2 for understanding temporal integration in black-box machine learning models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e697120d-6c87-4805-9490-093adf856e7cCited by top-tier papers3
- Large language models transition from integrating across position-yoked, exponential windows to structure-yoked, power-law windowsDavid Skrill, Samuel Norman-HaignereNeurIPS 2023 · 7 citations
- Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration CodingRui-Chen Zheng, Wenrui Liu, Hui-Peng Du, Qinglin Zhang et al.AAAI 2026 · 4 citations
- Information Integration in Large Language Models is Gated by Linguistic Structural MarkersWei Liu, Nai DingEMNLP 2025
Builds on1
Related papers
- A deep convolutional neural network that is invariant to time rescalingBrandon G. Jacques, Zoran Tiganj, Aakash Sarkar, Marc W. Howard et al.ICML 2022 · 10 citations
- Interpretable multi-timescale models for predicting fMRI responses to continuous natural speechShailee Jain, Vy A. Vo, Shivangi Mahto, Amanda LeBel et al.NeurIPS 2020 · 58 citations
- DeepSITH: Efficient Learning via Decomposition of What and When Across Time ScalesBrandon G. Jacques, Zoran Tiganj, Marc W. Howard, Per B. SederbergNeurIPS 2021 · 9 citations
- Language Through a Prism: A Spectral Approach for Multiscale Language RepresentationsAlex Tamkin, Dan Jurafsky, Noah D. GoodmanNeurIPS 2020 · 72 citations
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab et al.NeurIPS 2021 · 1,280 citations
