Lune

ICML2024Top-tier venue

Modeling Language Tokens as Functionals of Semantic Fields

Zhengqi Pei, Anran Zhang, Shuhui Wang, Qingming Huang

2024Year
2Citations
1Top-tier citations

Abstract

Recent advances in natural language processing have relied heavily on using Transformer-based language models. However, Transformers often require large parameter sizes and model depth. Existing Transformer-free approaches using statespace models demonstrate superiority over Transformers, yet they still lack a neuro-biologically connection to the human brain. This paper proposes LasF , representing Language tokens as Functionals of semantic fields, to simulate the neuronal behaviors for better language modeling. The LasF module is equivalent to a nonlinear approximator tailored for sequential data. By replacing the final layers of pre-trained language models with the LasF module, we obtain LasF -based models. Experiments conducted for standard reading comprehension and questionanswering tasks demonstrate that the LasF -based models consistently improve accuracy with fewer parameters. Besides, we use CommonsenseQA's blind test set to evaluate a full-parameter tuned LasF -based model, which outperforms the prior best ensemble and single models by 0.4% and 3.1%, respectively. Furthermore, our LasF -only language model trained from scratch outperforms existing parameter-efficient language models on standard datasets such as WikiText103 and Pen-nTreebank.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers1

Ask how each one uses it

Builds on10

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines