Language Through a Prism: A Spectral Approach for Multiscale Language Representations
Alex Tamkin, Dan Jurafsky, Noah D. Goodman
摘要
Language exhibits structure at different scales, ranging from subwords to words, sentences, paragraphs, and documents. To what extent do deep models capture information at these scales, and can we force them to better capture structure across this hierarchy? We approach this question by focusing on individual neurons, analyzing the behavior of their activations at different timescales. We show that signal processing provides a natural framework for separating structure across scales, enabling us to 1) disentangle scale-specific information in existing embeddings and 2) train models to learn more about particular scales. Concretely, we apply spectral filters to the activations of a neuron across an input, producing filtered embeddings that perform well on part of speech tagging (word-level), dialog speech acts classification (utterance-level), or topic classification (document-level), while performing poorly on the other tasks. We also present a prism layer for training models, which uses spectral filters to constrain different neurons to model structure at different scales. Our proposed BERT + Prism model can better predict masked tokens using long-range context and produces multiscale representations that perform better at utterance- and document-level tasks. Our methods are general and readily applicable to other domains besides language, such as images, audio, and video.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Frequency Enhanced Hybrid Attention Network for Sequential RecommendationXinyu Du, Huanhuan Yuan, Pengpeng Zhao, Jianfeng Qu 等SIGIR 2023 · 被引用 142 次
- Latent Fourier TransformMason Wang, Cheng-Zhi Anna HuangICLR 2026 · 被引用 58 次
- FreeDyG: Frequency Enhanced Continuous-Time Dynamic Graph Model for Link PredictionYuxing Tian, Yiyan Qi, Fan GuoICLR 2024 · 被引用 58 次
- Codebook Features: Sparse and Discrete Interpretability for Neural NetworksAlex Tamkin, Mohammad Taufeeque, Noah D. GoodmanICML 2024 · 被引用 42 次
- Language modeling via stochastic processesRose E. Wang, Esin Durmus, Noah D. Goodman, Tatsunori HashimotoICLR 2022 · 被引用 28 次
相关 Paper
- Interpretable multi-timescale models for predicting fMRI responses to continuous natural speechShailee Jain, Vy A. Vo, Shivangi Mahto, Amanda LeBel 等NeurIPS 2020 · 被引用 58 次
- Multi-Scale Self-Attention for Text ClassificationQipeng Guo, Xipeng Qiu, Pengfei Liu, Xiangyang Xue 等AAAI 2020 · 被引用 69 次
- Analyzing Individual Neurons in Pre-trained Language ModelsNadir Durrani, Hassan Sajjad, Fahim Dalvi, Yonatan BelinkovEMNLP 2020 · 被引用 5 次
- BrainBERT: Self-supervised representation learning for intracranial recordingsChristopher Wang, Vighnesh Subramaniam, Adam Uri Yaari, Gabriel Kreiman 等ICLR 2023 · 被引用 13 次
- Learning Multiscale Transformer Models for Sequence GenerationBei Li, Tong Zheng, Yi Jing, Chengbo Jiao 等ICML 2022 · 被引用 15 次
