ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
Cihang Peng, Qiming Hou, Zhong Ren, Kun Zhou
Abstract
We present ROVI, a high-quality synthetic dataset for instance-grounded text-to-image generation, created by labeling 1M curated web images. Our key innovation is a strategy called re-captioning, focusing on the pre-detection stage, where a VLM (Vision-Language Model) generates comprehensive visual descriptions that are then processed by an LLM (Large Language Model) to extract a flat list of potential categories for OVDs (Open-Vocabulary Detectors) to detect. This approach yields a global prompt inherently linked to instance annotations while capturing secondary visual elements humans typically overlook. Evaluations show that ROVI exceeds existing detection datasets in image quality and resolution while containing two orders of magnitude more categories with an open-vocabulary nature. For demonstrative purposes, a text-to-image model GLIGEN trained on ROVI significantly outperforms state-of-the-art alternatives in instance grounding accuracy, prompt fidelity, and aesthetic quality. Our dataset and reproducible pipeline are available at https://github.com/CihangPeng/ROVI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 70f143e8-4e9e-4a80-9bfd-fb6b16747e58Cited by top-tier papers1
Ask how each one uses itBuilds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language ModelsShenghao Fu, Qize Yang, Qijie Mo, Junkai Yan et al.CVPR 2025
- Multi-Modal Classifiers for Open-Vocabulary Object DetectionPrannay Kaul, Weidi Xie, Andrew ZissermanICML 2023 · 69 citations
- Hyperbolic Learning with Synthetic Captions for Open-World DetectionFanjie Kong, Yanbei Chen, Jiarui Cai, Davide ModoloCVPR 2024
- ROD-MLLM: Towards More Reliable Object Detection in Multimodal Large Language ModelsHeng Yin, Yuqiang Ren, Ke Yan, Shouhong Ding et al.CVPR 2025
- CapDet: Unifying Dense Captioning and Open-World Detection PretrainingYanxin Long, Youpeng Wen, Jianhua Han, Hang Xu et al.CVPR 2023
