FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model
Chunyu Xie, Bin Wang, Fanjing Kong, Jincheng Li, Dawei Liang, Ji Ao, Dawei Leng, Yuhui Yin
摘要
Fine-grained vision-language understanding requires precise alignment between visual content and linguistic descriptions, a capability that remains limited in current models, particularly in non-English settings. While models like CLIP perform well on global alignment, they often struggle to capture fine-grained details in object attributes, spatial relations, and linguistic expressions, with limited support for bilingual comprehension. To address these challenges, we introduce FG-CLIP 2, a bilingual vision-language model designed to advance fine-grained alignment for both English and Chinese. Our approach leverages rich fine-grained supervision, including region-text matching and long-caption modeling, alongside multiple discriminative objectives. We further introduce the Textual Intra-modal Contrastive (TIC) loss to better distinguish semantically similar captions. Trained on a carefully curated mixture of large-scale English and Chinese data, including a newly released 12M Chinese region-text dataset, FG-CLIP 2 achieves powerful bilingual performance. To enable rigorous evaluation, we present a new benchmark for Chinese multimodal understanding, featuring long-caption retrieval and bounding box classification. Extensive experiments on 29 datasets across 8 tasks show that FG-CLIP 2 outperforms existing methods, achieving state-of-the-art results in both languages. We release the model, code, and benchmark to facilitate future research on bilingual fine-grained vision-language alignment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- WeDetect: Fast Open-Vocabulary Object Detection as RetrievalShenghao Fu, Yukun Su, Fengyun Rao, Jing Lyu 等CVPR 2026 · 被引用 11 次
- LoCoT2V-Bench: Benchmarking Long-Form and Complex Text-to-Video GenerationXiangqing Zheng, CHENGYUE WU, Kehai Chen, Min zhangICML 2026 · 被引用 3 次
- From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific LiteratureKun Yuan, Min Woo Sun, Zhen Chen, Alejandro Lozano 等CVPR 2026
- ObjEmbed: Towards Universal Multimodal Object EmbeddingsShenghao Fu, Yukun Su, Fengyun Rao, Jing LYU 等ICML 2026
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 被引用 767 次
相关 Paper
- FG-CLIP: Fine-Grained Visual and Textual AlignmentChunyu Xie, Bin Wang, Fanjing Kong, Jincheng Li 等ICML 2025
- β-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language AlignmentFatimah Zohra, Chen Zhao, Hani Itani, Bernard GhanemCVPR 2026 · 被引用 6 次
- PixCLIP: Towards Fine-grained Vision-Language Understanding via Any-granularity Pixel-Text AlignmentYicheng Xiao, Yu Chen, Hao-Xuan Ma, Jiale Hong 等ICML 2026 · 被引用 4 次
- FLAIR: VLM with Fine-grained Language-informed Image RepresentationsRui Xiao, Sanghwan Kim, Mariana-Iuliana Georgescu, Zeynep Akata 等CVPR 2025
- Delving into Multimodal Prompting for Fine-Grained Visual ClassificationXin Jiang, Hao Tang, Junyao Gao, Xiaoyu Du 等AAAI 2024 · 被引用 71 次
