Understanding Adaptive, Multiscale Temporal Integration In Deep Speech Recognition Systems
Menoua Keshishian, Samuel Norman-Haignere, Nima Mesgarani
摘要
Natural signals such as speech are hierarchically structured across many different timescales, spanning tens (e.g., phonemes) to hundreds (e.g., words) of milliseconds, each of which is highly variable and context-dependent. While deep neural networks (DNNs) excel at recognizing complex patterns from natural signals, relatively little is known about how DNNs flexibly integrate across multiple timescales. Here, we show how a recently developed method for studying temporal integration in biological neural systems – the temporal context invariance (TCI) paradigm – can be used to understand temporal integration in DNNs. The method is simple: we measure responses to a large number of stimulus segments presented in two different contexts and estimate the smallest segment duration needed to achieve a context invariant response. We applied our method to understand how the popular DeepSpeech2 model learns to integrate across time in speech. We find that nearly all of the model units, even in recurrent layers, have a compact integration window within which stimuli substantially alter the response and outside of which stimuli have little effect. We show that training causes these integration windows to shrink at early layers and expand at higher layers, creating a hierarchy of integration windows across the network. Moreover, by measuring integration windows for time-stretched/compressed speech, we reveal a transition point, midway through the trained network, where integration windows become yoked to the duration of stimulus structures (e.g., phonemes or words) rather than absolute time. Similar phenomena were observed in a purely recurrent and purely convolutional network although structure-yoked integration was more prominent in the recurrent network. These findings suggest that deep speech recognition systems use a common motif to encode the hierarchical structure of speech: integrating across short, time-yoked windows at early layers and long, structure-yoked windows at later layers. Our method provides a straightforward and general-purpose toolkit2 for understanding temporal integration in black-box machine learning models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Large language models transition from integrating across position-yoked, exponential windows to structure-yoked, power-law windowsDavid Skrill, Samuel Norman-HaignereNeurIPS 2023 · 被引用 7 次
- Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration CodingRui-Chen Zheng, Wenrui Liu, Hui-Peng Du, Qinglin Zhang 等AAAI 2026 · 被引用 4 次
- Information Integration in Large Language Models is Gated by Linguistic Structural MarkersWei Liu, Nai DingEMNLP 2025
它引用的顶会 Paper1
相关 Paper
- A deep convolutional neural network that is invariant to time rescalingBrandon G. Jacques, Zoran Tiganj, Aakash Sarkar, Marc W. Howard 等ICML 2022 · 被引用 10 次
- Interpretable multi-timescale models for predicting fMRI responses to continuous natural speechShailee Jain, Vy A. Vo, Shivangi Mahto, Amanda LeBel 等NeurIPS 2020 · 被引用 58 次
- DeepSITH: Efficient Learning via Decomposition of What and When Across Time ScalesBrandon G. Jacques, Zoran Tiganj, Marc W. Howard, Per B. SederbergNeurIPS 2021 · 被引用 9 次
- Language Through a Prism: A Spectral Approach for Multiscale Language RepresentationsAlex Tamkin, Dan Jurafsky, Noah D. GoodmanNeurIPS 2020 · 被引用 72 次
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab 等NeurIPS 2021 · 被引用 1,280 次
