Lune

ICLR2026顶会

Adaptive Logit Adjustment for Debiasing Multimodal Language Models

Hoin Jung, Junyi Chai, Xiaoqian Wang

出版方
2026年份

摘要

Vision-Language Models (VLMs) and Large Multimodal Models (LMMs) have significantly advanced image-to-text generation tasks such as image captioning and visual question answering (VQA). However, these models often exhibit biases, including attribute misalignment between the generated text and the input image, or the reinforcement of harmful stereotypes. Existing debiasing techniques primarily focus on modifying representations at the encoder or decoder level, which can degrade model performance and may be susceptible to bias reintroduction from external sources. In this work, we propose Adaptive Logit Adjustment (ALA) for Bias Alignment and Neutralization, a post-hoc debiasing method that operates directly on logits during autoregressive text generation. Unlike prior approaches that modify internal representations, ALA selectively adjusts token probabilities to mitigate biases without distorting essential model outputs. Our approach leverages external classifiers to measure bias misalignment between image and text, applies gradient-based importance analysis to identify bias-inducing tokens, and dynamically refines token probabilities to reduce undesired biases. We evaluate ALA on image captioning and various VQA tasks, demonstrating its effectiveness in mitigating bias while maintaining contextual accuracy. Notably, our approach is applicable to various multimodal architectures in a model-agnostic manner, including VLMs and LMMs, across different tasks that involve autoregressive text generation. Our results show that logit-based debiasing offers a flexible and efficient alternative to existing encoder-and embedding-centric approaches, providing a more practical solution for building fairer multimodal AI systems. The code is available on GitHub. * Corresponding author. RELATED WORK BIAS IN IMAGE-TO-TEXT GENERATION Image captioning and VQA involve generating textual descriptions for images. Prior studies (Fraser & Kiritchenko, 2024; Sathe et al., 2024; Howard et al., 2024b;a; Girrbach et al., 2025) have highlighted the presence of bias in such image-to-text tasks as detailed in Section 3. While these studies effectively quantify biases in model outputs, most remain limited to observational analysis and do not propose

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 2369535f-1d72-4911-bbc6-e5164f584b72

它引用的顶会 Paper15

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖