From Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data Calibration
Mingyang Song, Xiaoye Qu, Jiawei Zhou, Yu Cheng
Abstract
Large Vision-Language Models (LVLMs) have achieved significant progress in combining visual comprehension with language generation.Despite this success, the training data of LVLMs still suffers from Long-Tail (LT) problems, where the data distribution is highly imbalanced.Previous works have mainly focused on traditional VLM architectures, i.e., CLIP or ViT, and specific tasks such as recognition and classification. Nevertheless, the exploration of LVLM (e.g. LLaVA) and more general tasks (e.g. Visual Question Answering and Visual Reasoning) remains under-explored.In this paper, we first conduct an in-depth analysis of the LT issues in LVLMs and identify two core causes: the overrepresentation of head concepts and the underrepresentation of tail concepts.Based on the above observation, we propose an Adaptive Data Refinement Framework (ADR), which consists of two stages: Data Rebalancing (DR) and Data Synthesis (DS).In the DR stage, we adaptively rebalance the redundant data based on entity distributions, while in the DS stage, we leverage Denoising Diffusion Probabilistic Models (DDPMs) and scarce images to supplement underrepresented portions.Through comprehensive evaluations across eleven benchmarks, our proposed ADR effectively mitigates the long-tail problem in the training data, improving the average performance of LLaVA 1.5 relatively by 4.36%, without increasing the training data volume. Our code and data will be publicly released.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Enhancing Vision-Language Model Reliability with Uncertainty-Guided Dropout DecodingYixiong Fang, Ziran Yang, Zhaorun Chen, Zhuokai Zhao et al.NeurIPS 2025 · 21 citations
- Reframing Long-Tailed Learning via Loss Landscape GeometryShenghan Chen, Yiming Liu, Yanzhen Wang, Yujia Wang et al.CVPR 2026 · 2 citations
- DIVER: Diving Deeper into Distilled Data via Expressive Semantic RecoveryQianxin Xia, Zhiyong Shu, Wenbo Jiang, Jiawei Du et al.ICML 2026
- Adaptive Token Refinement in Long-Tailed Large Vision-Language Models Fine-TuningWenjun Miao, Mingda Li, Yanchao Hao, Zheng WeiICML 2026
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Large Language Models Know What is Key Visual Entity: An LLM-assisted Multimodal Retrieval for VQAPu Jian, Donglei Yu, Jiajun ZhangEMNLP 2024 · 5 citations
- Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMsKanchana Ranasinghe, Satya Narayan Shukla, Omid Poursaeed, Michael S. Ryoo et al.CVPR 2024 · 21 citations
- Dual-Phase Visual-Language Pretraining and Adaptation for Long-Tailed Multi-Label RecognitionYongcheng Li, Xuekuan Wang, Zhifei Zhang, Cairong ZhaoAAAI 2026
- ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and WisdomJingqi Zhou, Sheng Wang, Jingwei Dong, Kai Liu et al.EMNLP 2025
- VideoPro: Adaptive Program Reasoning for Long Video UnderstandingChenglin Li, Feng Han, Yikun Wang, Ruilin Li et al.ACL 2026 · 4 citations
