FLASH: Fast Generative Retrieval via Autoregressive Semantic Hashing with Provably Distance Bounds
Yifei Zhang, Hao Zhu, Haoran Shi, Yanyu Chen, Xiaolin Han, Yupei Zhang, Lingyun Song, Wenxuan Wang, Chao Zhou, Xuequn Shang, Piotr Koniusz
Abstract
Retrieval-Augmented Generation (RAG) relies critically on the effectiveness of its retrieval component. Dense retrieval with Approximate Nearest Neighbor (ANN) search is efficient but relies on symmetric similarity metrics that can limit the modeling of directional relevance. Generative retrieval addresses this limitation by ranking documents through conditional generation, but existing approaches assign documents unstructured identifiers that lack geometric grounding, leading to error codes and poorly controlled ranking behavior. We propose Flash (Fast Generative Retrieval via Autoregressive Semantic Hashing), a framework that integrates random orthogonal projection hashing (ROPH) with autoregressive hash code generation. Documents are encoded into D-bit binary codes that provide two key properties: (i) a structured Hamming space in which imperfectly generated codes can still retrieve semantically related candidates, and (ii) an unbiased inner-product estimator with O(1/√D) error, enabling binary codes to support accurate similarity estimation. We introduce a dual-head autoregressive decoder that jointly predicts hash tokens and a quantization scalar, allowing the estimator to be applied directly to generated candidates while modeling asymmetric relevance through conditional generation. At inference time, candidate scoring and index storage reduce to bitwise operations, substantially improving efficiency compared to dense retrieval.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 33429c43-d1bf-4d51-976a-55f0774372f0Related papers
- Collision to Cognition: Hash-Driven Graph Construction for Efficient RAGChuang Zhou, Zheng Yuan, Linhao Luo, Zhaozhuo Xu et al.ACL 2026
- DReX: Accurate and Scalable Dense Retrieval Acceleration via Algorithmic-Hardware CodesignDerrick Quinn, E. Ezgi Yücel, Martin Prammer, Zhenxing Fan et al.ISCA 2025 · 10 citations
- Unsupervised Multi-Index Semantic HashingChristian Hansen, Casper Hansen, Jakob Grue Simonsen, Stephen Alstrup et al.WWW 2021 · 11 citations
- Content-aware Neural Hashing for Cold-start RecommendationCasper Hansen, Christian Hansen, Jakob Grue Simonsen, Stephen Alstrup et al.SIGIR 2020 · 29 citations
- Planning Ahead in Generative Retrieval: Guiding Autoregressive Generation through Simultaneous DecodingHansi Zeng, Chen Luo, Hamed ZamaniSIGIR 2024 · 21 citations
