IVQA-LD: Inclusive Multimodal Understanding for Population with Limb Deficiency
Yan Ke, Xin Shen, Jiaying Ying, Xin Li, Xin Yu
Abstract
People with limb differences often face significant challenges in accessing inclusive AI services, largely due to the lack of structured, high-quality resources centered on disability contexts. In this work, we introduce a limb-deficiency aware body-centric learning and evaluation paradigm that involves (i) a large-scale limb-aware vision-language dataset and evaluation benchmark for multimodal reasoning, and (ii) a model adaptation strategy for Vision-Language Models (VLM) in limb-difference contexts. Specifically, we first collect limb-difference data covering all eight limb-deficiency types across diverse real-world scenarios. The data are systematically organized into 96 limb-affected human action categories and 68 functional classes derived from internationally recognized classification frameworks. Then, we curate a vision-language dataset incorporating expert annotations for limb-aware multimodal understanding, named Inclusive VQA for Limb Deficiency (IVQA-LD). IVQA-LD comprises 80K VQA pairs spanning eight core tasks including visual grounding, quantitative reasoning, functional semantic classification, and instructional text generation. We benchmark state-of-the-art VLMs on IVQA-LD and find that they struggle across all tasks, exposing substantial deficiencies in limb-aware perception and reasoning. To address this, we further propose a Bodycentric Structure-aware Initialization (BSI) strategy that aligns model representations with limbspecific semantics. With BSI, VLMs fine-tuned on IVQA-LD achieve significant performance improvements across all the tasks. We publicly release the dataset to support future research. The code and data are available at IVQA-LD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 52c9b2a1-3be4-4e89-9555-c79e37413c84Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- WalkVLM: Aid Visually Impaired People Walking by Vision Language ModelZhiqiang Yuan, Ting Zhang, Yeshuang Zhu, Jiapei Zhang et al.ICCV 2025 · 3 citations
- Right this way: Can VLMs Guide Us to See More to Answer Questions?Li Liu, Diji Yang, Sijia Zhong, Kalyana Suma Sree Tholeti et al.NeurIPS 2024 · 20 citations
- PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical EnvironmentsWeijie Zhou, Xuantang Xiong, Yi Peng, Manli Tao et al.NeurIPS 2025 · 4 citations
- OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLMYutao Hu, Tianbin Li, Quanfeng Lu, Wenqi Shao et al.CVPR 2024
- LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of LivingDominick Reilly, Rajatsubhra Chakraborty, Arkaprava Sinha, Manish Kumar Govind et al.CVPR 2025
