FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity
Hang Hua, Qing Liu, Lingzhi Zhang, Jing Shi, Soo Ye Kim, Zhifei Zhang, Yilin Wang, Jianming Zhang, Zhe Lin, Jiebo Luo
Abstract
In this sunset-lit beach scene captured from an eye-level view, a wooden table holds a green glass vase with brown and white patterns, containing orange and pink flowers with black cores and green leaves. The vase casts a shadow on the table, which also holds a glass of orange liquid, a bowl of colorful fruits, and a neatly arranged plate of food with utensils beside it. A red wine glass sits on the far right, balancing the composition. The objects are arranged harmoniously, with their glossy surfaces reflecting the warm light, creating a peaceful, tropical dining setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Latent Chain-of-Thought for Visual ReasoningGuohao Sun, Hang Hua, Jian Wang, Jiebo Luo et al.NeurIPS 2025 · 30 citations
- Top-Down Semantic Refinement for Image CaptioningJusheng Zhang, Kaitong Cai, Jing Yang, Jian Wang et al.AAAI 2026 · 16 citations
- On the Generalization Capacities of MLLMs for Spatial IntelligenceGongjie Zhang, Wenhao Li, Quanhao Qian, Jiuniu Wang et al.ICLR 2026 · 10 citations
- ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart UnderstandingJovana Kondic, Pengyuan Li, Dhiraj Joshi, Isaac Sanchez et al.CVPR 2026 · 7 citations
- ExCap3d: Expressive 3D Scene Understanding via Object Captioning with Varying DetailChandan Yeshwanth, Dávid Rozenberszki, Angela DaiICCV 2025 · 2 citations
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- WMarkGPT: Watermarked Image Understanding via Multimodal Large Language ModelsSongbai Tan, Xuerui Qiu, Yao Shu, Gang Xu et al.ICML 2025
- UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in RLRui Tian, Mingfei Gao, Haiming Gang, Jiasen Lu et al.CVPR 2026
- EditAR: Unified Conditional Generation with Autoregressive ModelsJiteng Mu, Nuno Vasconcelos, Xiaolong WangCVPR 2025
- FullDiT: Video Generative Foundation Models with Multimodal Control via Full AttentionXuan Ju, Weicai Ye, Quande Liu, Qiulin Wang et al.ICCV 2025 · 5 citations
- Language-driven Grasp DetectionVuong Dinh An, Minh Nhat Vu, Baoru Huang, Nghia Nguyen et al.CVPR 2024
