O3SLM: Open Weight, Open Data, and Open Vocabulary Sketch-Language Model
Rishi Gupta, Mukilan Karuppasamy, Shyam Marjit, Aditay Tripathi, Anirban Chakraborty
摘要
While Large Vision Language Models (LVLMs) are increasingly deployed in real-world applications, their ability to interpret abstract visual inputs remains limited. Specifically, they struggle to comprehend hand-drawn sketches, a modality that offers an intuitive means of expressing concepts that are difficult to describe textually. We identify the primary bottleneck as the absence of a large-scale dataset that jointly models sketches, photorealistic images, and corresponding natural language instructions. To address this, we present two key contributions: (1) a new, large-scale dataset of image-sketch-instruction triplets designed to facilitate both pretraining and instruction tuning, and (2) O3SLM, an LVLM trained on this dataset. Comprehensive evaluations on multiple sketch-based tasks: (a) object localization, (b) counting, (c) image retrieval i.e., (SBIR and fine-grained SBIR), and (d) visual question answering (VQA); while incorporating the three existing sketch datasets, namely QuickDraw!, Sketchy, and Tu-Berlin, along with our generated SketchVCL dataset, show that O3SLM achieves state-of-the-art performance, substantially outperforming existing LVLMs in sketch comprehension and reasoning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Objects365: A Large-Scale, High-Quality Dataset for Object DetectionShuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng 等ICCV 2019 · 被引用 1,018 次
- SpatialRGPT: Grounded Spatial Reasoning in Vision-Language ModelsAn-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo 等NeurIPS 2024 · 被引用 412 次
- CLIPasso: semantically-aware object sketchingYael Vinker, Ehsan Pajouheshgar, Jessica Y. Bo, Roman Christian Bachmann 等SIGGRAPH 2022 · 被引用 219 次
相关 Paper
- SketchMind: Understanding Abstract Sketches with MLLMs for Fine-Grained Sketch-Based Image RetrievalChangxing Li, Donglin Zhang, Zhikai Hu, Xiao-Jun Wu 等WWW 2026
- BOP-ASK: Object-Interaction Reasoning for Vision-Language ModelsVineet Bhat, Sungsu Kim, Valts Blukis, Greg Heinrich 等CVPR 2026 · 被引用 6 次
- OURO: A Self-Bootstrapped Framework for Enhancing Multimodal Scene UnderstandingTianrun Xu, Guanyu Chen, Ye Li, Yuxin Xi 等ICCV 2025 · 被引用 2 次
- MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing UnderstandingQian Kou, Xiaofeng Shi, Yulin Li, Xiaosong Qiu 等ICML 2026 · 被引用 2 次
- Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language ModelsLei Li, Yuqi Wang, Runxin Xu, Peiyi Wang 等ACL 2024 · 被引用 16 次
