Ultra-High Dimensional Sparse Representations with Binarization for Efficient Text Retrieval
Kyoungrok Jang, Junmo Kang, Giwon Hong, Sung-Hyon Myaeng, Joohee Park, Taewon Yoon, Hee-Cheol Seo
Abstract
The semantic matching capabilities of neural information retrieval can ameliorate synonymy and polysemy problems of symbolic approaches. However, neural models' dense representations are more suitable for re-ranking, due to their inefficiency. Sparse representations, either in symbolic or latent form, are more efficient with an inverted index. Taking the merits of the sparse and dense representations, we propose an ultra-high dimensional (UHD) representation scheme equipped with directly controllable sparsity. UHD's large capacity and minimal noise and interference among the dimensions allow for binarized representations, which are highly efficient for storage and search. Also proposed is a bucketing method, where the embeddings from multiple layers of BERT are selected/merged to represent diverse linguistic aspects. We test our models with MS MARCO and TREC CAR, showing that our models outperforms other sparse models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 373d9e2b-ee61-4e9c-8354-188bdf41bed3Cited by top-tier papers4
- Crafting Interpretable Embeddings for Language Neuroscience by Asking LLMs QuestionsVinamra Benara, Chandan Singh, John X. Morris, Richard J. Antonello et al.NeurIPS 2024 · 26 citations
- Efficient Neural Ranking using Forward IndexesJurek Leonhardt, Koustav Rudra, Megha Khosla, Abhijit Anand et al.WWW 2022 · 16 citations
- STAIR: Learning Sparse Text and Image Representation in Grounded TokensChen Chen, Bowen Zhang, Liangliang Cao, Jiguang Shen et al.EMNLP 2023 · 15 citations
- From Tokens to Concepts: Leveraging SAE for SPLADEYuxuan Zong, Mathias Vast, Basile Van Cooten, Laure Soulier et al.SIGIR 2026
Builds on5
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Context-Aware Document Term Weighting for Ad-Hoc SearchZhuyun Dai, Jamie CallanWWW 2020 · 123 citations
- Minimizing FLOPs to Learn Efficient Sparse RepresentationsBiswajit Paria, Chih-Kuan Yeh, Ian En-Hsu Yen, Ning Xu et al.ICLR 2020 · 85 citations
- Roles and Utilization of Attention Heads in Transformer-based Neural Language ModelsJae-young Jo, Sung-Hyon MyaengACL 2020 · 32 citations
- SOLAR: Sparse Orthogonal Learned and Random EmbeddingsTharun Medini, Beidi Chen, Anshumali ShrivastavaICLR 2021 · 10 citations
Related papers
- Adaptive Sparsity Optimization with Learnable Soft Top-K and Per-Term Thresholding for Efficient RetrievalWentai Xie, Parker Carlson, Shanxiu He, Tao YangSIGIR 2026
- SDR: Efficient Neural Re-ranking using Succinct Document RepresentationNachshon Cohen, Amit Portnoy, Besnik Fetahu, Amir IngberACL 2022
- Million-Scale Text-to-Video Retrieval with Hyperdimensional ComputingHyunsei Lee, Jaewoo Gwak, Shinhyoung Jang, Junyoung Lee et al.EuroSys 2026
- PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document RetrievalShengyao Zhuang, Xueguang Ma, Bevan Koopman, Jimmy Lin et al.EMNLP 2024 · 26 citations
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
