F-LMM: Grounding Frozen Large Multimodal Models
Size Wu, Sheng Jin, Wenwei Zhang, Lumin Xu, Wentao Liu, Wei Li, Chen Change Loy
Abstract
Figure 1. An example of user-AI conversation around an image. Left: The current state-of-the-art grounding model GLaMM [60] is effective for grounded conversation when prompted by "answer with interleaved masks", but fails to follow user instruction to answer a single word (yes or no) and misunderstands the question as a referring segmentation prompt. Right: Our F-LMM preserves instructionfollowing ability while being able to perform visual grounding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08e487d5-1830-45db-bc18-393b1e5e4e6aCited by top-tier papers8
- WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian NavigationRafi Ibn Sultan, Hui Zhu, Xiangyu Zhou, Chengyin Li et al.CVPR 2026 · 4 citations
- GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual GroundingPeirong Zhang, Yidan Zhang, Luxiao Xu, Jinliang Lin et al.CVPR 2026 · 3 citations
- Harmonizing Visual Representations for Unified Multimodal Understanding and GenerationSize Wu, Wenwei Zhang, Lumin Xu, Sheng Jin et al.ICCV 2025 · 3 citations
- iRULER: Intelligible Rubric-Based User-Defined LLM Evaluation for RevisionJingwen Bai, Wei Soon Cheong, Philippe Muller, Brian Y. LimCHI 2026 · 1 citation
- Instruction-guided Multi-Granularity Segmentation and Captioning with Large Multimodal ModelXu Yuan, Li Zhou, Zenghui Sun, Zikun Zhou et al.AAAI 2025 · 1 citation
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
Related papers
- GLaMM: Pixel Grounding Large Multimodal ModelHanoona Abdul Rasheed, Muhammad Maaz, Sahal Shaji Mullappilly, Abdelrahman M. Shaker et al.CVPR 2024 · 113 citations
- VideoGLaMM : A Large Multimodal Model for Pixel-Level Visual Grounding in VideosShehan Munasinghe, Hanan Gani, Wenqi Zhu, Jiale Cao et al.CVPR 2025
- Instruction-Guided Visual MaskingJinliang Zheng, Jianxiong Li, Sijie Cheng, Yinan Zheng et al.NeurIPS 2024 · 22 citations
- SegLLM: Multi-round Reasoning Segmentation with Large Language ModelsXudong Wang, Shaolun Zhang, Shufan Li, Kehan Li et al.ICLR 2025
- GeoPixel: Pixel Grounding Large Multimodal Model in Remote SensingAkashah Shabbir, Mohammed Zumri, Mohammed Bennamoun, Fahad Shahbaz Khan et al.ICML 2025
