Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken Generation
Yongqi Li, Hongru Cai, Wenjie Wang, Leigang Qu, Yinwei Wei, Wenjie Li, Liqiang Nie, Tat-Seng Chua
摘要
Text-to-image retrieval is a fundamental task in multimedia retrieval.Traditional studies have typically approached this task as a discriminative problem, matching the text and image via the cross-attention mechanism (one-tower framework) or in a common embedding space (two-tower framework).The one-tower framework excels in effectiveness but falls short in efficiency, whereas the two-tower framework is efficient but struggles to maintain competitive effectiveness.In this study, we aim to enhance both effectiveness and efficiency by transforming the text-to-image retrieval task into a token-to-voken generation problem, where fine-grained interactions are incorporated to improve effectiveness while maintaining high efficiency.Despite its potential advantages, this paradigm shift presents significant challenges: 1) misalignment with high-level semantics and 2) learning gap towards the retrieval target.To address the challenges, we propose AVG, which discretizes images into vokens while aligning with both the visual information and high-level semantics.Additionally, to bridge the learning gap between generative training and the retrieval target, AVG incorporates discriminative training to modify the learning direction during token-to-voken training.Experiments demonstrate that the benefits of paradigm innovation are realized: compared with the classical two-tower method, CLIP, AVG achieves the 7.53% relative effectiveness improvement and also 4× efficiency improvement.We release code at the GitHub repository.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Pretrain-then-Adapt: Uncertainty-Aware Test-Time Adaptation for Text-based Person SearchJiahao Zhang, Shaofei Huang, Yaxiong Wang, Zhedong ZhengSIGIR 2026
- GENIUS: A Generative Framework for Universal Multimodal SearchSungyeon Kim, Xinliang Zhu, Xiaofan Lin, Muhammet Bastan 等CVPR 2025
- CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented GenerationKaiwen Wei, Xiao Liu, Jie Zhang, Zijian Wang 等WWW 2026
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Vector-quantized Image Modeling with Improved VQGANJiahui Yu, Xin Li, Jing Yu Koh, Han Zhang 等ICLR 2022 · 被引用 753 次
相关 Paper
- TempMe: Video Temporal Token Merging for Efficient Text-Video RetrievalLeqi Shen, Tianxiang Hao, Tao He, Sicheng Zhao 等ICLR 2025
- COTS: Collaborative Two-Stream Vision-Language Pre-Training Model for Cross-Modal RetrievalHaoyu Lu, Nanyi Fei, Yuqi Huo, Yizhao Gao 等CVPR 2022 · 被引用 68 次
- Concept-Guided Tokenization: Closing the Gap Between Reconstruction and GenerationYunqiao Yang, Haokun Lin, Guanzhong Wu, Ying WeiICML 2026
- Vokenization: Improving Language Understanding with Contextualized, Visual-Grounded SupervisionHao Tan, Mohit BansalEMNLP 2020 · 被引用 72 次
- Hybrid-Tower: Fine-Grained Pseudo-Query Interaction and Generation for Text-to-Video RetrievalBangxiang Lan, Ruobing Xie, Ruixiang Zhao, Xingwu Sun 等ICCV 2025 · 被引用 5 次
