Tuning Computer Vision Models With Task Rewards
André Susano Pinto, Alexander Kolesnikov, Yuge Shi, Lucas Beyer, Xiaohua Zhai
Abstract
Misalignment between model predictions and intended usage can be detrimental for the deployment of computer vision models. The issue is exacerbated when the task involves complex structured outputs, as it becomes harder to design procedures which address this misalignment. In natural language processing, this is often addressed using reinforcement learning techniques that align models with a task reward. We adopt this approach and show its surprising effectiveness across multiple computer vision tasks, such as object detection, panoptic segmentation, colorization and image captioning. We believe this approach has the potential to be widely useful for better aligning models with a diverse range of computer vision tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 65f56bfb-129d-4ba0-9a9a-c4b13837856fCited by top-tier papers21
- Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewardsAlexandre Ramé, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya et al.NeurIPS 2023 · 295 citations
- 4M: Massively Multimodal Masked ModelingDavid Mizrahi, Roman Bachmann, Oguzhan Fatih Kar, Teresa Yeo et al.NeurIPS 2023 · 154 citations
- LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis EvaluationYujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang et al.NeurIPS 2023 · 119 citations
- Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language ModelsShuai Zhao, Xiaohan Wang, Linchao Zhu, Yi YangICLR 2024 · 47 citations
- LocCa: Visual Pretraining with Location-aware CaptionersBo Wan, Michael Tschannen, Yongqin Xian, Filip Pavetic et al.NeurIPS 2024 · 38 citations
Builds on11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Objects365: A Large-Scale, High-Quality Dataset for Object DetectionShuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng et al.ICCV 2019 · 1,018 citations
Related papers
- CapWAP: Image Captioning with a PurposeAdam Fisch, Kenton Lee, Ming-Wei Chang, Jonathan H. Clark et al.EMNLP 2020 · 17 citations
- A Unified Sequence Interface for Vision TasksTing Chen, Saurabh Saxena, Lala Li, Tsung-Yi Lin et al.NeurIPS 2022 · 201 citations
- CCCaption: Dual-Reward Reinforcement Learning for Complete and Correct Image CaptioningZhijiang Tang, Linhua Wang, Jiaxin Qi, Weihao Jiang et al.CVPR 2026 · 7 citations
- Mirage or Method? How Model–Task Alignment Induces Divergent RL ConclusionsHaoze Wu, Cheng Wang, Wenshuo Zhao, Junxian HeICLR 2026 · 7 citations
- UViM: A Unified Modeling Approach for Vision with Learned Guiding CodesAlexander Kolesnikov, André Susano Pinto, Lucas Beyer, Xiaohua Zhai et al.NeurIPS 2022 · 88 citations
