FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity
Hang Hua, Qing Liu, Lingzhi Zhang, Jing Shi, Soo Ye Kim, Zhifei Zhang, Yilin Wang, Jianming Zhang, Zhe Lin, Jiebo Luo
摘要
In this sunset-lit beach scene captured from an eye-level view, a wooden table holds a green glass vase with brown and white patterns, containing orange and pink flowers with black cores and green leaves. The vase casts a shadow on the table, which also holds a glass of orange liquid, a bowl of colorful fruits, and a neatly arranged plate of food with utensils beside it. A red wine glass sits on the far right, balancing the composition. The objects are arranged harmoniously, with their glossy surfaces reflecting the warm light, creating a peaceful, tropical dining setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Latent Chain-of-Thought for Visual ReasoningGuohao Sun, Hang Hua, Jian Wang, Jiebo Luo 等NeurIPS 2025 · 被引用 30 次
- Top-Down Semantic Refinement for Image CaptioningJusheng Zhang, Kaitong Cai, Jing Yang, Jian Wang 等AAAI 2026 · 被引用 16 次
- On the Generalization Capacities of MLLMs for Spatial IntelligenceGongjie Zhang, Wenhao Li, Quanhao Qian, Jiuniu Wang 等ICLR 2026 · 被引用 10 次
- ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart UnderstandingJovana Kondic, Pengyuan Li, Dhiraj Joshi, Isaac Sanchez 等CVPR 2026 · 被引用 7 次
- ExCap3d: Expressive 3D Scene Understanding via Object Captioning with Varying DetailChandan Yeshwanth, Dávid Rozenberszki, Angela DaiICCV 2025 · 被引用 2 次
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- WMarkGPT: Watermarked Image Understanding via Multimodal Large Language ModelsSongbai Tan, Xuerui Qiu, Yao Shu, Gang Xu 等ICML 2025
- UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in RLRui Tian, Mingfei Gao, Haiming Gang, Jiasen Lu 等CVPR 2026
- EditAR: Unified Conditional Generation with Autoregressive ModelsJiteng Mu, Nuno Vasconcelos, Xiaolong WangCVPR 2025
- FullDiT: Video Generative Foundation Models with Multimodal Control via Full AttentionXuan Ju, Weicai Ye, Quande Liu, Qiulin Wang 等ICCV 2025 · 被引用 5 次
- Language-driven Grasp DetectionVuong Dinh An, Minh Nhat Vu, Baoru Huang, Nghia Nguyen 等CVPR 2024
