X-SAM: From Segment Anything to Any Segmentation
Hao Wang, Limeng Qiao, Zequn Jie, Zhijian Huang, Chengjian Feng, Qingfang Zheng, Lin Ma, Xiangyuan Lan, Xiaodan Liang
Abstract
Large Language Models (LLMs) demonstrate strong capabilities in broad knowledge representation, yet they are inherently deficient in pixel-level perceptual understanding. Although the Segment Anything Model (SAM) represents a significant advancement in visual-prompt-driven image segmentation, it exhibits notable limitations in multi-mask prediction and category-specific segmentation tasks, and it cannot integrate all segmentation tasks within a unified model architecture. To address these limitations, we present X-SAM, a streamlined Multimodal Large Language Model (MLLM) framework that extends the segmentation paradigm from segment anything to any segmentation. Specifically, we introduce a novel unified framework that enables more advanced pixel-level perceptual comprehension for MLLMs. Furthermore, we propose a new segmentation task, termed Visual GrounDed (VGD) segmentation, which segments all instance objects with interactive visual prompts and empowers MLLMs with visual grounded, pixel-wise interpretative capabilities. To enable effective training on diverse data sources, we present a unified training strategy that supports co-training across multiple datasets. Experimental results demonstrate that X-SAM achieves state-of-the-art performance on a wide range of image segmentation benchmarks, highlighting its efficiency for multimodal, pixel-level visual understanding. Code is available at https://github.com/ wanghao9610/X-SAM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cbd603b9-2509-42ee-bcf3-365789e9dc87Cited by top-tier papers7
- SAM 3: Segment Anything with ConceptsNicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath et al.ICLR 2026 · 1,103 citations
- VoxTell: Free-Text Promptable Universal 3D Medical Image SegmentationMaximilian Rokuss, Moritz Langenberg, Yannick Kirchhoff, Fabian Isensee et al.CVPR 2026 · 22 citations
- SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing ImagesZepeng Xin, Kaiyu Li, Luodi Chen, Wanchen Li et al.CVPR 2026 · 14 citations
- SAMTok: Representing Any Mask with Two WordsYikang Zhou, Tao Zhang, Dengxian Gong, Yuanzheng Wu et al.CVPR 2026 · 10 citations
- SegCompass: Exploring Interpretable Alignment with Sparse Autoencoders for Enhanced Reasoning SegmentationZhenyu Lu, Liupeng Li, Jinpeng Wang, Haoqian Kang et al.CVPR 2026 · 3 citations
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
Related papers
- GLaMM: Pixel Grounding Large Multimodal ModelHanoona Abdul Rasheed, Muhammad Maaz, Sahal Shaji Mullappilly, Abdelrahman M. Shaker et al.CVPR 2024 · 113 citations
- OpenWorldSAM: Extending SAM2 for Universal Image Segmentation with Language PromptsShiting Xiao, Rishabh Kabra, Yuhang Li, Donghyun Lee et al.NeurIPS 2025 · 15 citations
- Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and VideosWeifeng Lin, Xinyu Wei, Ruichuan An, Tianhe Ren et al.NeurIPS 2025 · 47 citations
- Hugging Visual Prompt and Segmentation Tokens: Consistency Learning for Fine-Grained Visual Understanding in MLLMsjing yang, Sen Yang, Boqiang Duan, Ming Dai et al.CVPR 2026
- Multi-Modal Instruction Tuned LLMs with Fine-Grained Visual PerceptionJunwen He, Yifan Wang, Lijun Wang, Huchuan Lu et al.CVPR 2024
