Socialformer: Social Network Inspired Long Document Modeling for Document Ranking
Yujia Zhou, Zhicheng Dou, Huaying Yuan, Zhengyi Ma
Abstract
Utilizing pre-trained language models has achieved great success for neural document ranking. Limited by the computational and memory requirements, long document modeling becomes a critical issue. Recent works propose to modify the full attention matrix in Transformer by designing sparse attention patterns. However, most of them only focus on local connections of terms within a fixed-size window. How to build suitable remote connections between terms to better model document representation remains underexplored. In this paper, we propose the model Socialformer, which introduces the characteristics of social networks into designing sparse attention patterns for long document modeling in document ranking. Specifically, we consider several attention patterns to construct a graph like social networks. Endowed with the characteristic of social networks, most pairs of nodes in such a graph can reach with a short path while ensuring the sparsity. To facilitate efficient calculation, we segment the graph into multiple subgraphs to simulate friend circles in social scenarios. Experimental results confirm the effectiveness of our model on long document modeling. CCS CONCEPTS • Information systems → Retrieval models and ranking.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc348367-b838-4367-a3ec-2465cd44148aCited by top-tier papers2
- Webformer: Pre-training with Web Pages for Information RetrievalYu Guo, Zhengyi Ma, Jiaxin Mao, Hongjin Qian et al.SIGIR 2022 · 30 citations
- A Silver Bullet or a Compromise for Full Attention? A Comprehensive Study of Gist Token-based Context CompressionChenlong Deng, Zhisong Zhang, Kelong Mao, Shuaiyi Li et al.ACL 2025
Builds on10
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang et al.AAAI 2021 · 7,289 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- ETC: Encoding Long and Structured Inputs in TransformersJoshua Ainslie, Santiago Ontañón, Chris Alberti, Vaclav Cvicek et al.EMNLP 2020 · 268 citations
Related papers
- Even Sparser Graph TransformersHamed Shirzad, Honghao Lin, Balaji Venkatachalam, Ameya Velingker et al.NeurIPS 2024 · 18 citations
- A Graph-based Relevance Matching Model for Ad-hoc RetrievalYufeng Zhang, Jinghao Zhang, Zeyu Cui, Shu Wu et al.AAAI 2021 · 26 citations
- LongRanker: Efficient One-Pass Document Reranking with Long-Context Large Language ModelsChangjiang Zhou, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke et al.WWW 2026
- ERNIE-Doc: A Retrospective Long-Document Modeling TransformerSiyu Ding, Junyuan Shang, Shuohuan Wang, Yu Sun et al.ACL 2021
- WebFormer: The Web-page Transformer for Structure Information ExtractionQifan Wang, Yi Fang, Anirudh Ravula, Fuli Feng et al.WWW 2022 · 88 citations
