Adaptive Token Refinement in Long-Tailed Large Vision-Language Models Fine-Tuning
Wenjun Miao, Mingda Li, Yanchao Hao, Zheng Wei
摘要
While large vision-language models (LVLMs) have shown remarkable adaptability to downstream applications, their fine-tuning process remains susceptible to bias under long-tailed data. Compared to zero-shot scenarios, fine-tuning LVLMs on imbalanced datasets often yields limited performance improvements on tail data. This is because LVLMs tend to rapidly overfit the head data at an early fine-tuning stage, thereby impairing the learning of the tail data while simultaneously failing to exploit their quantitative advantage. Furthermore, in many downstream LVLM scenarios, quantified long-tailed prior knowledge of data distribution is often unavailable, significantly limiting the applicability of traditional long-tailed techniques that rely heavily on such information. To address these issues, we propose the Adaptive Token Refinement (ATR), a novel framework that adaptively refines the learning process of LVLMs under long-tailed data. Specifically, ATR consists of two token-level operations applied to output and input tokens, respectively: 1) a bounded adaptive loss that dynamically filters and reweights output tokens to mitigate overfitting on head data, and 2) a visual token mask strategy that augments the probability paths of input tokens to enhance long-tailed performance. Extensive experiments demonstrate that ATR consistently enhance both performance and generalization for long-tailed LVLMs fine-tuning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan 等ICLR 2020 · 被引用 1,496 次
- Long-tail learning via logit adjustmentAditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain 等ICLR 2021 · 被引用 937 次
相关 Paper
- From Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data CalibrationMingyang Song, Xiaoye Qu, Jiawei Zhou, Yu ChengCVPR 2025
- Improving CLIP Adaptation by Breaking Tail Alignment for Source-Free Cross-Domain Few-Shot LearningShuai Yi, Yixiong Zou, Yuhua Li, Ruixuan LiICML 2026 · 被引用 1 次
- Uniformly Distributed Category Prototype-Guided Vision-Language Framework for Long-Tail RecognitionXiaoxuan He, Siming Fu, Xinpeng Ding, Yuchen Cao 等ACM MM 2023 · 被引用 6 次
- LTGC: Long-Tail Recognition via Leveraging LLMs-Driven Generated ContentQihao Zhao, Yalun Dai, Hao Li, Wei Hu 等CVPR 2024 · 被引用 22 次
- Bag of Tricks for Long-Tailed Visual Recognition with Deep Convolutional Neural NetworksYongshun Zhang, Xiu-Shen Wei, Boyan Zhou, Jianxin WuAAAI 2021 · 被引用 162 次
