ICLR2024

Democratizing Fine-grained Visual Recognition with Large Language Models

Mingxuan Liu, Subhankar Roy, Wenjing Li, Zhun Zhong, Nicu Sebe, Elisa Ricci

被引用 27 次

摘要

Well. Could you describe this photo and its wing color, head pattern, …, primary color ? B: : I see a bird in a photo. How to distinguish its specific species? Reasoning Concepts from Observations Inference … Reasoning For Each Sample : Certainly. The bird is perched on a tree branch amidst the falling snow. Its wings are grey, and it boasts a black and red pattern on its head. Notably, its dominant color is red. : Perfect. Even though I can't see it, but based on your description, I think the bird you see would be a Pyrrhuloxia, Cardinal, or Summer Tanager. Few Unlabeled Observations … VLM Test Images Semantic Classification with Reasoned Concepts Gadwall Cardinal Red Eyed Vireo Rusty Blackbird Lincoln Sparrow Tropical Kingbird … … Large Language Model Visual Question Answering Model Vision-Langauge Model Figure 1: An overview of our proposed fine-grained visual recognition (FGVR) pipeline. Left: Given few unlabelled images we exploit visual question answering (VQA) and large language models (LLM) to reason about subordinate-level category names without requiring expert knowledge. Right: At inference, we utilize the reasoned concepts to carry out FGVR via zero-shot semantic classification with a vision-language model (VLM).