Tag-grounded Visual Instruction Tuning with Retrieval Augmentation
Daiqing Qi, Handong Zhao, Zijun Wei, Sheng Li
Abstract
GPT4:This photo features four mangosteens on a wooden table.Three of them are uncut, showing their round, dark purple exterior.The fourth mangosteen is cut open, revealing its white, fleshy interior, which is divided into segments.The table surface has a natural wood grain pattern and adds a rustic touch to the image. LLaVA:The image features a close-up of three small, round, and purple fruits, possibly pomegranates, sitting on a wooden surface.The fruits are placed in a row, with one on the left, one in the middle, and one on the right.In addition to the fruits, there is a knife located on the right side of the image, likely used for cutting the fruits.The scene appears to be a still life composition, showcasing the vibrant color and texture of the fruits. ShareGPT4V:The image captures a close-up view of four lychee fruits, their vibrant colors standing out against the wooden surface they rest on.The fruits are arranged in a square formation, with one at each corner.The top left and bottom right corners each hold a whole lychee, their green tops contrasting with the deep purple of the fruit.The top right corner features a lychee that has been cut open, revealing the white interior that is characteristic of this tropical fruit.The background, though blurred, hints at the presence of green leaves, adding a touch of nature to the scene.What type of fruit is this?
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bdbc9fe8-a1bf-4d44-a34e-1897548695b1Cited by top-tier papers2
- Improve Temporal Reasoning in Multimodal Large Language Models via Video Contrastive DecodingDaiqing Qi, Dongliang Guo, Hanzhang Yuan, Handong Zhao et al.NeurIPS 2025 · 5 citations
- The Photographer's Eye: Teaching Multimodal Large Language Models to See, and Critique Like PhotographersDaiqing Qi, Handong Zhao, Jing Shi, Simon Jenni et al.CVPR 2025
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
Related papers
- Wi-Fruit: See Through Fruits with Smart DevicesYutong Liu, Landu Jiang, Linghe Kong, Qiao Xiang et al.UbiComp 2022 · 39 citations
- FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any GranularityHang Hua, Qing Liu, Lingzhi Zhang, Jing Shi et al.CVPR 2025
- Metric Transforms and Low Rank Representations of Kernels for Fast AttentionTimothy Chu, Josh Alman, Gary L. Miller, Shyam Narayanan et al.NeurIPS 2024 · 4 citations
- VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame SelectionSonghao Han, Wei Huang, Hairong Shi, Le Zhuo et al.CVPR 2025
- ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion ModelsOzgur Kara, Krishna Kumar Singh, Feng Liu, Duygu Ceylan et al.CVPR 2025
