Lune

ICML2026顶会

IVQA-LD: Inclusive Multimodal Understanding for Population with Limb Deficiency

Yan Ke, Xin Shen, Jiaying Ying, Xin Li, Xin Yu

出版方
2026年份

摘要

People with limb differences often face significant challenges in accessing inclusive AI services, largely due to the lack of structured, high-quality resources centered on disability contexts. In this work, we introduce a limb-deficiency aware body-centric learning and evaluation paradigm that involves (i) a large-scale limb-aware vision-language dataset and evaluation benchmark for multimodal reasoning, and (ii) a model adaptation strategy for Vision-Language Models (VLM) in limb-difference contexts. Specifically, we first collect limb-difference data covering all eight limb-deficiency types across diverse real-world scenarios. The data are systematically organized into 96 limb-affected human action categories and 68 functional classes derived from internationally recognized classification frameworks. Then, we curate a vision-language dataset incorporating expert annotations for limb-aware multimodal understanding, named Inclusive VQA for Limb Deficiency (IVQA-LD). IVQA-LD comprises 80K VQA pairs spanning eight core tasks including visual grounding, quantitative reasoning, functional semantic classification, and instructional text generation. We benchmark state-of-the-art VLMs on IVQA-LD and find that they struggle across all tasks, exposing substantial deficiencies in limb-aware perception and reasoning. To address this, we further propose a Bodycentric Structure-aware Initialization (BSI) strategy that aligns model representations with limbspecific semantics. With BSI, VLMs fine-tuned on IVQA-LD achieve significant performance improvements across all the tasks. We publicly release the dataset to support future research. The code and data are available at IVQA-LD.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper14

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖