RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm
Tiancheng Gu, Kaicheng Yang, Chaoyi Zhang, Yin Xie, Xiang An, Ziyong Feng, Dongnan Liu, Weidong Cai, Jiankang Deng
摘要
After pre-training on extensive image-text pairs, Contrastive Language-Image Pre-training (CLIP) demonstrates promising performance on a wide variety of benchmarks. However, a substantial volume of multimodal interleaved documents remains underutilized for contrastive vision-language representation learning. To fully leverage these unpaired documents, we initially establish a Real-World Data Extraction pipeline to extract high-quality images and texts. Then we design a hierarchical retrieval method to efficiently associate each image with multiple semantically relevant realistic texts. To further enhance fine-grained visual information, we propose an image semantic augmented generation module for synthetic text production. Furthermore, we employ a semantic balance sampling strategy to improve dataset diversity, enabling better learning of long-tail concepts. Based on these innovations, we construct RealSyn, a dataset combining realistic and synthetic texts, available in three scales: 15M, 30M, and 100M. We compare our dataset with other widely used datasets of equivalent scale for CLIP training. Models pre-trained on RealSyn consistently achieve state-of-the-art performance across various downstream tasks, including linear probe, zero-shot transfer, zero-shot robustness, and zero-shot retrieval. Furthermore, extensive experiments confirm that RealSyn significantly enhances contrastive vision-language representation learning and demonstrates robust scalability. The code will be released in https://garygutc.github.io/RealSyn.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Reverse-Engineered Reasoning for Open-Ended GenerationHaozhe Wang, Haoran Que, Qixin Xu, Minghao Liu 等ICLR 2026 · 被引用 38 次
- UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding LearningTiancheng Gu, Kaicheng Yang, Kaichen Zhang, Xiang An 等AAAI 2026 · 被引用 24 次
- Decoupled Global-Local Alignment for Improving Compositional UnderstandingXiaoxing Hu, Kaicheng Yang, Jun Wang, Haoran Xu 等ACM MM 2025 · 被引用 4 次
- Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person RetrievalTianlu Zheng, Yifan Zhang, Xiang An, Ziyong Feng 等EMNLP 2025 · 被引用 1 次
它引用的顶会 Paper33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin 等ICML 2022 · 被引用 1,058 次
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesDouwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami 等NeurIPS 2020 · 被引用 1,022 次
相关 Paper
- ALIP: Adaptive Language-Image Pre-training with Synthetic CaptionKaicheng Yang, Jiankang Deng, Xiang An, Jiawei Li 等ICCV 2023 · 被引用 93 次
- RWKV-CLIP: A Robust Vision-Language Representation LearnerTiancheng Gu, Kaicheng Yang, Xiang An, Ziyong Feng 等EMNLP 2024 · 被引用 11 次
- TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language NegativesMaitreya Patel, Abhiram Kusumba, Sheng Cheng, Changhoon Kim 等NeurIPS 2024 · 被引用 73 次
- Non-Contrastive Learning Meets Language-Image Pre-TrainingJinghao Zhou, Li Dong, Zhe Gan, Lijuan Wang 等CVPR 2023
- RA-CLIP: Retrieval Augmented Contrastive Language-Image Pre-TrainingChen-Wei Xie, Siyang Sun, Xiong Xiong, Yun Zheng 等CVPR 2023
