Temporal Precision Matters: Brain-Tuning Speech Language Models with Millisecond-Resolution Neural Signals
Zhejun Zhang, Wenqing Zhou, Haozhe Xu, Lin Zhang, Lei Li
摘要
Brain-tuning enhances brain alignment and downstream performance by fine-tuning speech language models with neural recordings. However, previous work relies primarily on fMRI, whose temporal resolution integrates neural activity over seconds, blending distinct processing stages into a single supervision signal and precluding temporally targeted training. We introduce ECoG-tuning, which leverages electro-corticography’s millisecond precision to train speech language models. We design temporally targeted windows—a speech window capturing acoustic-phonetic encoding and a language window capturing higher-order linguistic processing—grounded in neuroscientific findings about temporal encoding hierarchies. Evaluating three models on the Podcast ECoG dataset, we find that ECoG-tuning significantly improves brain alignment over pretrained and distillation baselines. Notably, full spatiotemporal dynamics yield 7–17% higher alignment than time-averaged supervision across models, and language-window tuning produces larger gains in higher-order language regions, indicating that temporal precision provides additional training value. Moreover, ECoG-tuned models consistently improve or maintain down-stream performance. Overall, our work provides initial evidence that electrophysiology is a viable brain-tuning modality, demonstrating how neuroscientific insights into processing hierarchies can inform principled model training strategies. Code is available at https://github. com/Mochizuki-BUPT/ECoG-Tuning-main.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Toward a realistic model of speech processing in the brain with self-supervised learningJuliette Millet, Charlotte Caucheteux, Pierre Orhan, Yves Boubenec 等NeurIPS 2022 · 被引用 164 次
- Brain-Informed Fine-Tuning for Improved Multilingual Understanding in Language ModelsAnuja Negi, Subba Reddy Oota, Anwar Nunez-Elizalde, Manish Gupta 等NeurIPS 2025 · 被引用 9 次
- Brain-tuning Improves Generalizability and Efficiency of Brain Alignment in Speech ModelsOmer Moussa, Mariya TonevaNeurIPS 2025 · 被引用 7 次
相关 Paper
- Improving Semantic Understanding in Speech Language Models via Brain-tuningOmer Moussa, Dietrich Klakow, Mariya TonevaICLR 2025
- Abstraction Induces the Brain Alignment of Language and Speech ModelsEmily Cheng, Aditya Vaidya, Richard AntonelloICML 2026
- Spontaneous Yet Predictable: Shapelet-Driven, Channel-Aware Intention Decoding from Multi-Region ECoGKeren Cao, Yuhang Tian, Kaizhong Zheng, Wei Xi 等AAAI 2026
- Language models and brains align due to more than next-word prediction and word-level informationGabriele Merlin, Mariya TonevaEMNLP 2024 · 被引用 2 次
- Multimodal Scaling Laws for Task & Data-Optimized Models of Visual CortexAbdülkadir Gökce, Yingtian Tang, Martin SchrimpfICML 2026
