Chunk, Align, Select: A Simple Long-sequence Processing Method for Transformers
Jiawen Xie, Pengyu Cheng, Xiao Liang, Yong Dai, Nan Du
2024Year
3Top-tier citations
Abstract
Although dominant in natural language processing, transformer-based models remain challenged by the task of long-sequence processing,
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 74ad433f-bc47-4c71-9608-06b393e7ce8eCited by top-tier papers3
- Self-playing Adversarial Language Game Enhances LLM ReasoningPengyu Cheng, Tianhao Hu, Han Xu, Zhisong Zhang et al.NeurIPS 2024 · 120 citations
- DocHieNet: A Large and Diverse Dataset for Document Hierarchy ParsingHangdi Xing, Changxu Cheng, Feiyu Gao, Zirui Shao et al.EMNLP 2024 · 3 citations
- Overcoming Vocabulary Mismatch: Vocabulary-agnostic Teacher Guided Language ModelingHaebin Shin, Lei Ji, Xiao Liu, Yeyun GongICML 2025
Builds on9
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 1,143 citations
- Memorizing TransformersYuhuai Wu, Markus Norman Rabe, DeLesley Hutchins, Christian SzegedyICLR 2022 · 231 citations
- PRIMERA: Pyramid-based Masked Sentence Pre-training for Multi-document SummarizationWen Xiao, Iz Beltagy, Giuseppe Carenini, Arman CohanACL 2022 · 147 citations
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song et al.ICLR 2021 · 122 citations
Related papers
- Hierarchical Walking Transformer for Object Re-IdentificationXudong Tian, Jun Liu, Zhizhong Zhang, Chengjie Wang et al.ACM MM 2022 · 6 citations
- The Unstoppable Rise of Computational Linguistics in Deep LearningJames HendersonACL 2020 · 4 citations
- Task Descriptors Help Transformers Learn Linear Models In-ContextRuomin Huang, Rong GeICLR 2025
- Dream Video: Composing Your Dream Videos with Customized Subject and MotionYujie Wei, Shiwei Zhang, Zhiwu Qing, Hangjie Yuan et al.CVPR 2024
- AnchorAttention: Difference-Aware Sparse Attention with Stripe GranularityYu Zhang, Dong Guo, Fang Wu, Guoliang Zhu et al.EMNLP 2025
