CALICO: Part-Focused Semantic Co-Segmentation with Large Vision-Language Models
Kiet A. Nguyen, Adheesh Sunil Juvekar, Tianjiao Yu, Muntasir Wahed, Ismini Lourentzou
Abstract
Can you segment the common object in these images? The common object is the dog. Both images show a dog. The common part is the head. The images show a dog. The unique parts are the neck (IMAGE1), the leg, the foot, the body, and the tail (IMAGE2). Common Parts: Please segment the common parts of the objects. Unique Parts: What are the unique parts of the objects in these images? Please output segmentation masks. CALICO Figure 1. Multi-Image Part-focused Object Comparison with CALICO. Our pixel-grounded Large Vision-Language Model, CALICO, performs part-focused semantic co-segmentation, a newly introduced task where the goal is to identify, segment, and label common objects, as well as common and unique object parts across multiple images.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47d5e474-d799-495c-9b77-fac67801eb66Cited by top-tier papers4
- Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMsYifan Shen, Yuanzhe Liu, Jingyuan Zhu, Xu Cao et al.NeurIPS 2025 · 41 citations
- MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMsYunqiu Xu, Linchao Zhu, Yi YangICCV 2025 · 7 citations
- LangHOPS: Language Grounded Hierarchical Open-Vocabulary Part SegmentationYang Miao, Jan-Nico Zaech, Xi Wang, Fabien Despinoy et al.NeurIPS 2025 · 3 citations
- Segment and Matte Anything in a Unified ModelZezhong Fan, Xiaohan Li, Topojoy Biswas, Kaushiki Nag et al.AAAI 2026
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- PACO: Parts and Attributes of Common ObjectsVignesh Ramanathan, Anmol Kalia, Vladan Petrovic, Yi Wen et al.CVPR 2023
- Unveiling Parts Beyond Objects: Towards Finer-Granularity Referring Expression SegmentationWenxuan Wang, Tongtian Yue, Yisi Zhang, Longteng Guo et al.CVPR 2024 · 5 citations
- GSVA: Generalized Segmentation via Multimodal Large Language ModelsZhuofan Xia, Dongchen Han, Yizeng Han, Xuran Pan et al.CVPR 2024 · 42 citations
- Pixel Aligned Language ModelsJiarui Xu, Xingyi Zhou, Shen Yan, Xiuye Gu et al.CVPR 2024 · 6 citations
- Groundhog Grounding Large Language Models to Holistic SegmentationYichi Zhang, Ziqiao Ma, Xiaofeng Gao, Suhaila Shakiah et al.CVPR 2024 · 24 citations
