SECRET: Towards Scalable and Efficient Code Retrieval via Segmented Deep Hashing
Wenchao Gu, Ensheng Shi, Yanlin Wang, Lun Du, Shi Han, Hongyu Zhang, Dongmei Zhang, Michael R. Lyu
摘要
Code retrieval, which retrieves code snippets based on users' natural language descriptions, is widely used by devel-opers and plays a pivotal role in real-world software development. The advent of deep learning has shifted the retrieval paradigm from lexical-based matching towards leveraging deep learning models to encode source code and queries into vector represen-tations, facilitating code retrieval according to vector similarity. Despite the effectiveness of these models, managing large-scale code database presents significant challenges. Previous research proposes deep hashing-based methods, which generate hash codes for queries and code snippets and use Hamming distance for rapid recall of code candidates. However, this approach's reliance on linear scanning of the entire code base limits its scalability. To further improve the efficiency of large-scale code retrieval, we propose a novel approach SECRET (Scalable and Efficient Code Retrieval via SegmEnTed deep hashing). SECRET converts long hash codes calculated by existing deep hashing approaches into several short hash code segments through an iterative training strategy. After training, SECRET recalls code candidates by looking up the hash tables for each segment, the time complexity of recall can thus be greatly reduced. Extensive experimental results demonstrate that SECRET can drastically reduce the retrieval time by at least 95 % while achieving comparable or even higher performance of existing deep hashing approaches. Besides, SECRET also exhibits superior performance and efficiency compared to the classical hash table-based approach known as LSH under the same number of hash tables.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- DrainCode: Stealthy Energy Consumption Attacks on Retrieval-Augmented Code Generation via Context PoisoningYanli Wang, Jiadong Wu, Tianyue Jiang, Mingwei Liu 等ASE 2025 · 被引用 3 次
- AlignCoder: Aligning Retrieval with Target Intent for Repository-Level Code CompletionTianyue Jiang, Yanlin Wang, Yanli Wang, Daya Guo 等ASE 2025 · 被引用 2 次
它引用的顶会 Paper9
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Deep Joint-Semantics Reconstructing Hashing for Large-Scale Unsupervised Cross-Modal RetrievalShupeng Su, Zhisheng Zhong, Chao ZhangICCV 2019 · 被引用 261 次
- Joint-modal Distribution-based Similarity Hashing for Large-scale Unsupervised Deep Cross-modal RetrievalSong Liu, Shengsheng Qian, Yang Guan, Jiawei Zhan 等SIGIR 2020 · 被引用 214 次
- SPT-Code: Sequence-to-Sequence Pre-Training for Learning Source Code RepresentationsChangan Niu, Chuanyi Li, Vincent Ng, Jidong Ge 等ICSE 2022 · 被引用 99 次
相关 Paper
- Accelerating Code Search with Deep Hashing and Code ClassificationWenchao Gu, Yanlin Wang, Lun Du, Hongyu Zhang 等ACL 2022
- Asymmetric Deep Hashing for Efficient Hash Code CompressionShu Zhao, Dayan Wu, Wanqian Zhang, Yu Zhou 等ACM MM 2020 · 被引用 18 次
- Accelerate Learning of Deep Hashing With Gradient AttentionLong-Kai Huang, Jianda Chen, Sinno Jialin PanICCV 2019 · 被引用 22 次
- Unleashing the Full Potential of Product Quantization for Large-Scale Image RetrievalYu Liang, Shiliang Zhang, Li Ken Li, Xiaoyu WangNeurIPS 2023 · 被引用 5 次
- GHashing: Semantic Graph Hashing for Approximate Similarity Search in Graph DatabasesZongyue Qin, Yunsheng Bai, Yizhou SunKDD 2020 · 被引用 29 次
