VisualHow: Multimodal Problem Solving
Jinhui Yang, Xianyu Chen, Ming Jiang, Shi Chen, Louis Wang, Qi Zhao
摘要
Recent progress in the interdisciplinary studies of computer vision (CV) and natural language processing (NLP) has enabled the development of intelligent systems that can describe what they see and answer questions accordingly. However, despite showing usefulness in performing these vision-language tasks, existing methods still struggle in understanding real-life problems (i.e., how to do something) and suggesting step-by-step guidance to solve them. With an overarching goal of developing intelligent systems to assist humans in various daily activities, we propose Vi-sualHow, a free-form and open-ended research that focuses on understanding a real-life problem and deriving its solution by incorporating key components across multiple modalities. We develop a new dataset with 20,028 real-life problems and 102,933 steps that constitute their solutions, where each step consists of both a visual illustration and a textual description that guide the problem solving. To establish better understanding of problems and solutions, we also provide annotations of multimodal attention that localizes important components across modalities and solution graphs that encapsulate different steps in structured representations. These data and annotations enable a family of new vision-language tasks that solve real-life problems. Through extensive experiments with representative models, we demonstrate their effectiveness on training and testing models for the new tasks, and there is significant scope for improvement by learning effective attention mechanisms. Our dataset and models are available at https://github.com/formidify/VisualHow . * Equal contributions. Create an ornament with their picture on it. Cut out previous photos of your pet and make a collage. Play with your animal around the holidays. Give your pet a gift.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Explainable Saliency: Articulating Reasoning with Contextual PrioritizationNuo Chen, Ming Jiang, Qi ZhaoCVPR 2025
- Beyond Average: Individualized Visual Scanpath PredictionXianyu Chen, Ming Jiang, Qi ZhaoCVPR 2024
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li 等ICCV 2019 · 被引用 598 次
- MERLOT: Multimodal Neural Script Knowledge ModelsRowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu 等NeurIPS 2021 · 被引用 463 次
- VD-BERT: A Unified Vision and Dialog Transformer with BERTYue Wang, Shafiq R. Joty, Michael R. Lyu, Irwin King 等EMNLP 2020 · 被引用 68 次
相关 Paper
- MULTISCRIPT: Multimodal Script Learning for Supporting Open Domain Everyday TasksJingyuan Qi, Minqian Liu, Ying Shen, Zhiyang Xu 等AAAI 2024 · 被引用 3 次
- From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data SynthesisChuanqi Cheng, Jian Guan, Wei Wu, Rui YanEMNLP 2024 · 被引用 1 次
- Pretrained Language Models as Visual Planners for Human AssistanceDhruvesh Patel, Hamid Eghbalzadeh, Nitin Kamra, Michael Louis Iuzzolino 等ICCV 2023 · 被引用 41 次
- VisionMath: Vision-Form Mathematical Problem-SolvingZongyang Ma, Yuxin Chen, Ziqi Zhang, Zhongang Oi 等ICCV 2025 · 被引用 2 次
- WalkVLM: Aid Visually Impaired People Walking by Vision Language ModelZhiqiang Yuan, Ting Zhang, Yeshuang Zhu, Jiapei Zhang 等ICCV 2025 · 被引用 3 次
