Modeling Language Tokens as Functionals of Semantic Fields
Zhengqi Pei, Anran Zhang, Shuhui Wang, Qingming Huang
摘要
Recent advances in natural language processing have relied heavily on using Transformer-based language models. However, Transformers often require large parameter sizes and model depth. Existing Transformer-free approaches using statespace models demonstrate superiority over Transformers, yet they still lack a neuro-biologically connection to the human brain. This paper proposes LasF , representing Language tokens as Functionals of semantic fields, to simulate the neuronal behaviors for better language modeling. The LasF module is equivalent to a nonlinear approximator tailored for sequential data. By replacing the final layers of pre-trained language models with the LasF module, we obtain LasF -based models. Experiments conducted for standard reading comprehension and questionanswering tasks demonstrate that the LasF -based models consistently improve accuracy with fewer parameters. Besides, we use CommonsenseQA's blind test set to evaluate a full-parameter tuned LasF -based model, which outperforms the prior best ensemble and single models by 0.4% and 3.1%, respectively. Furthermore, our LasF -only language model trained from scratch outperforms existing parameter-efficient language models on standard datasets such as WikiText103 and Pen-nTreebank.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- Hyena Hierarchy: Towards Larger Convolutional Language ModelsMichael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu 等ICML 2023 · 被引用 481 次
相关 Paper
- Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language ProcessingZihang Dai, Guokun Lai, Yiming Yang, Quoc LeNeurIPS 2020 · 被引用 273 次
- State Space Models are Provably Comparable to Transformers in Dynamic Token SelectionNaoki Nishikawa, Taiji SuzukiICLR 2025
- State-Free Inference of State-Space Models: The Transfer Function ApproachRom N. Parnichkun, Stefano Massaroli, Alessandro Moro, Jimmy T. H. Smith 等ICML 2024 · 被引用 18 次
- Implicit Language Models are RNNs: Balancing Parallelization and ExpressivityMark Schöne, Babak Rahmani, Heiner Kremer, Fabian Falck 等ICML 2025
- DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted AveragingMatteo Pagliardini, Amirkeivan Mohtashami, François Fleuret, Martin JaggiNeurIPS 2024 · 被引用 60 次
