F-LMM: Grounding Frozen Large Multimodal Models
Size Wu, Sheng Jin, Wenwei Zhang, Lumin Xu, Wentao Liu, Wei Li, Chen Change Loy
2025年份
8顶会引用
摘要
Figure 1. An example of user-AI conversation around an image. Left: The current state-of-the-art grounding model GLaMM [60] is effective for grounded conversation when prompted by "answer with interleaved masks", but fails to follow user instruction to answer a single word (yes or no) and misunderstands the question as a referring segmentation prompt. Right: Our F-LMM preserves instructionfollowing ability while being able to perform visual grounding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian NavigationRafi Ibn Sultan, Hui Zhu, Xiangyu Zhou, Chengyin Li 等CVPR 2026 · 被引用 4 次
- GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual GroundingPeirong Zhang, Yidan Zhang, Luxiao Xu, Jinliang Lin 等CVPR 2026 · 被引用 3 次
- Harmonizing Visual Representations for Unified Multimodal Understanding and GenerationSize Wu, Wenwei Zhang, Lumin Xu, Sheng Jin 等ICCV 2025 · 被引用 3 次
- iRULER: Intelligible Rubric-Based User-Defined LLM Evaluation for RevisionJingwen Bai, Wei Soon Cheong, Philippe Muller, Brian Y. LimCHI 2026 · 被引用 1 次
- Instruction-guided Multi-Granularity Segmentation and Captioning with Large Multimodal ModelXu Yuan, Li Zhou, Zenghui Sun, Zikun Zhou 等AAAI 2025 · 被引用 1 次
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
相关 Paper
- GLaMM: Pixel Grounding Large Multimodal ModelHanoona Abdul Rasheed, Muhammad Maaz, Sahal Shaji Mullappilly, Abdelrahman M. Shaker 等CVPR 2024 · 被引用 113 次
- VideoGLaMM : A Large Multimodal Model for Pixel-Level Visual Grounding in VideosShehan Munasinghe, Hanan Gani, Wenqi Zhu, Jiale Cao 等CVPR 2025
- Instruction-Guided Visual MaskingJinliang Zheng, Jianxiong Li, Sijie Cheng, Yinan Zheng 等NeurIPS 2024 · 被引用 22 次
- SegLLM: Multi-round Reasoning Segmentation with Large Language ModelsXudong Wang, Shaolun Zhang, Shufan Li, Kehan Li 等ICLR 2025
- GeoPixel: Pixel Grounding Large Multimodal Model in Remote SensingAkashah Shabbir, Mohammed Zumri, Mohammed Bennamoun, Fahad Shahbaz Khan 等ICML 2025
