The World of an Octopus: How Reporting Bias Influences a Language Model's Perception of Color
Cory Paik, Stéphane Aroca-Ouellette, Alessandro Roncone, Katharina Kann
Abstract
Recent work has raised concerns about the inherent limitations of text-only pretraining. In this paper, we first demonstrate that reporting bias, the tendency of people to not state the obvious, is one of the causes of this limitation, and then investigate to what extent multimodal training can mitigate this issue. To accomplish this, we 1) generate the Color Dataset (CoDa), a dataset of human-perceived color distributions for 521 common objects; 2) use CoDa to analyze and compare the color distribution found in text, the distribution captured by language models, and a human's perception of color; and 3) investigate the performance differences between text-only and multimodal models on CoDa. Our results show that the distribution of colors that a language model recovers correlates more strongly with the inaccurate distribution found in text than with the ground-truth, supporting the claim that reporting bias negatively impacts and inherently limits text-only training. We then demonstrate that multimodal models can leverage their visual training to mitigate these effects, providing a promising avenue for future research. * *Email has no accent, but includes the hyphen. Everyone knows that most bananas are [MASK].
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- DesCo: Learning Object Recognition with Rich Language DescriptionsLiunian Harold Li, Zi-Yi Dou, Nanyun Peng, Kai-Wei ChangNeurIPS 2023 · 38 citations
- Linearly Mapping from Image to Text SpaceJack Merullo, Louis Castricato, Carsten Eickhoff, Ellie PavlickICLR 2023 · 25 citations
- Mind's Eye: Grounded Language Model Reasoning through SimulationRuibo Liu, Jason Wei, Shixiang Shane Gu, Te-Yen Wu et al.ICLR 2023 · 22 citations
- Mitigating Reporting Bias in Semi-supervised Temporal Commonsense Inference with Probabilistic Soft LogicBibo Cai, Xiao Ding, Bowen Chen, Li Du et al.AAAI 2022 · 12 citations
- Visually Grounded Commonsense Knowledge AcquisitionYuan Yao, Tianyu Yu, Ao Zhang, Mengdi Li et al.AAAI 2023 · 6 citations
Builds on8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Climbing towards NLU: On Meaning, Form, and Understanding in the Age of DataEmily M. Bender, Alexander KollerACL 2020 · 914 citations
- Inducing Relational Knowledge from BERTZied Bouraoui, José Camacho-Collados, Steven SchockaertAAAI 2020 · 183 citations
Related papers
- Are Vision-Language Transformers Learning Multimodal Representations? A Probing PerspectiveEmmanuelle Salin, Badreddine Farah, Stéphane Ayache, Benoît FavreAAAI 2022 · 49 citations
- Visually-Augmented Language ModelingWeizhi Wang, Li Dong, Hao Cheng, Haoyu Song et al.ICLR 2023 · 5 citations
- Evaluating Model Perception of Color Illusions in Photorealistic ScenesLingjun Mao, Zineng Tang, Alane SuhrCVPR 2025
- Leveraging Visual Knowledge in Language Tasks: An Empirical Study on Intermediate Pre-training for Cross-Modal Knowledge TransferWoojeong Jin, Dong-Ho Lee, Chenguang Zhu, Jay Pujara et al.ACL 2022
- Measuring Text-Image Retrieval Fairness with Synthetic DataLluís GómezSIGIR 2025 · 1 citation
