Token-Context Attention for NLI: An Alternative to Self-Attention
Xin Zhang, Victor S. Sheng
Abstract
Despite the rapid progress in large language models (LLMs), even sub-billion-scale systems perform at chance level on challenging natural language inference (NLI) benchmarks such as Adversarial Natural Language Inference (ANLI), while training larger models is often impractical due to limited computational resources. We address this parameter-efficiency bottleneck in NLI with a Complex-Vector Token Representation that explicitly decouples each token from its context, and a Token-Context Attention mechanism that updates each token based on the most informative contextual semantics. On ANLI, a 0.8B-parameter Token-Context Attention model achieves higher parameter efficiency (accuracy per parameter) than all 1B and comparable 0.8B self-attention baselines; it also suffers smaller performance degradation under Fast Gradient Sign Method (FGSM) and Projected Gradient Descent (PGD) attacks and achieves the largest few-shot gains on SNLI and MNLI while exhibiting no significant degradation in ANLI accuracy after adaptation. These results suggest that explicitly disentangling token and context offers a viable alternative to standard self-attention for NLI tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f9202cd-9a6a-40fa-9fbc-9d6bed57ed96Builds on9
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- Gated Linear Attention Transformers with Hardware-Efficient TrainingSonglin Yang, Bailin Wang, Yikang Shen, Rameswar Panda et al.ICML 2024 · 390 citations
- InfoBERT: Improving Robustness of Language Models from An Information Theoretic PerspectiveBoxin Wang, Shuohang Wang, Yu Cheng, Zhe Gan et al.ICLR 2021 · 132 citations
- NLP From Scratch Without Large-Scale Pretraining: A Simple and Efficient FrameworkXingcheng Yao, Yanan Zheng, Xiaocong Yang, Zhilin YangICML 2022 · 50 citations
- EulerFormer: Sequential User Behavior Modeling with Complex Vector AttentionZhen Tian, Wayne Xin Zhao, Changwang Zhang, Xin Zhao et al.SIGIR 2024 · 8 citations
Related papers
- A Unified Sparse Attention via Multi-Granularity CompressionSiran Liu, Zheng Cao, Yongchao HeICML 2026
- Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token SelectionDongwon Jo, Beomseok Kang, Jiwon Song, jae-joon kimICML 2026 · 1 citation
- Toward Adversarial Training on Contextualized Language RepresentationHongqiu Wu, Yongxiang Liu, Hanwen Shi, Hai Zhao et al.ICLR 2023 · 4 citations
- AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and PruningYiwu Zhong, Zhuoming Liu, Yin Li, Liwei WangICCV 2025 · 1 citation
- Latent-Condensed Transformer for Efficient Long Context ModelingZeng You, Yaofo Chen, Qiuwu Chen, Ying Sun et al.ACL 2026
