Modeling Collaborator: Enabling Subjective Vision Classification with Minimal Human Effort via LLM Tool-Use
Imad Eddine Toubal, Aditya Avinash, Neil Gordon Alldrin, Jan Dlabal, Wenlei Zhou, Enming Luo, Otilia Stretcu, Hao Xiong, Chun-Ta Lu, Howard Zhou, Ranjay Krishna, Ariel Fuxman, Tom Duerig
摘要
From content moderation to wildlife conservation, the number of applications that require models to recognize nuanced or subjective visual concepts is growing. Traditionally, developing classifiers for such concepts requires substantial manual effort measured in hours, days, or even months to identify and annotate data needed for training. Even with recently proposed Agile Modeling techniques, which enable rapid bootstrapping of image classifiers, users are still required to spend 30 minutes or more of monotonous, repetitive data labeling just to train a single classifier. Drawing on Fiske's Cognitive Miser theory, we propose a new framework that alleviates manual effort by replacing human labeling with natural language interactions, reducing the total effort required to define a concept by an order of magnitude: from labeling 2,000 images to only 100 plus some natural language interactions. Our framework leverages recent advances in foundation models, both large language models and vision-language models, to carve out the concept space through conversation and by automatically labeling training data points. Most importantly, our framework eliminates the need for crowdsourced annotations. Moreover, our framework ultimately produces lightweight classification models that are deployable in cost-sensitive scenarios. Across 15 subjective concepts and across 2 public image classification datasets, our trained models outperform traditional Agile Modeling as well as state-of-the-art zero-shot classification models like ALIGN, CLIP, CuPL, and large visual question answering models like PaLI-X. * This work was done during an internship at Google. † This work was done during working at Google. Modeling Collaborator (Ours) Decision: Positive Reasons: The image shows a plate of food with chopsticks on it. The food is a salad with tuna, pineapple, avocado, and sesame seeds. The salad is on a plate with chopsticks on it. The dish is gourmet and contains tuna. PaLI-X (baseline) Prompt: Is this tuna gourmet? Justify your answer. PaLI: No, this tuna is not gourmet. It is a simple salad with tuna, sesame seeds, and vegetables. It is not a fancy or elaborate dish. Prompt: Could this be an image of gourmet tuna? PaLI: Yes, this could be an image of gourmet tuna. Prompt: Is this tuna gourmet? Are you sure? PaLI: I'm not sure if this tuna is gourmet. It looks like it could be, but it's hard to say for sure without tasting it.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering TaskSunqi Fan, Jiashuo Cui, Meng-Hao Guo, Shuojin YangNeurIPS 2025 · 被引用 15 次
- Promptimizer: User-Led Prompt Optimization for Personal Content ClassificationLeijie Wang, Kathryn Yurechko, Amy X. ZhangCHI 2026 · 被引用 1 次
- Agile Deliberation: Concept Deliberation for Subjective Visual ClassificationLeijie Wang, Otilia Stretcu, Wei Qiao, Thomas Denby 等CVPR 2026
- Semantic and Expressive Variations in Image Captions Across LanguagesAndre Ye, Sebastin Santy, Jena D. Hwang, Amy X. Zhang 等CVPR 2025
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
相关 Paper
- Agile Modeling: From Concept to Classifier in MinutesOtilia Stretcu, Edward Vendrow, Kenji Hata, Krishnamurthy Viswanathan 等ICCV 2023 · 被引用 19 次
- Visual Classification via Description from Large Language ModelsSachit Menon, Carl VondrickICLR 2023 · 被引用 57 次
- Language-Driven Cross-Modal Classifier for Zero-Shot Multi-Label Image RecognitionYicheng Liu, Jie Wen, Chengliang Liu, Xiaozhao Fang 等ICML 2024 · 被引用 7 次
- Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image ClassificationYue Yang, Artemis Panagopoulou, Shenghao Zhou, Daniel Jin 等CVPR 2023
- ChatGPT-Powered Hierarchical Comparisons for Image ClassificationZhiyuan Ren, Yiyang Su, Xiaoming LiuNeurIPS 2023 · 被引用 54 次
