X-World: Accessibility, Vision, and Autonomy Meet
Jimuyang Zhang, Minglan Zheng, Matthew Boyd, Eshed Ohn-Bar
Abstract
An important issue facing vision-based intelligent systems today is the lack of accessibility-aware development. A main reason for this issue is the absence of any large-scale, standardized vision benchmarks that incorporate relevant tasks and scenarios related to people with disabilities. This lack of representation hinders even preliminary analysis with respect to underlying pose, appearance, and occlusion characteristics of diverse pedestrians. What is the impact of significant occlusion from a wheelchair on instance segmentation quality? How can interaction with mobility aids, e.g., a long and narrow walking cane, be recognized robustly? To begin addressing such questions, we introduce X-World, an accessibility-centered development environment for vision-based autonomous systems. We tackle inherent data scarcity by leveraging a simulation environment to spawn dynamic agents with various mobility aids. The simulation supports generation of ample amounts of finely annotated, multi-modal data in a safe, cheap, and privacy-preserving manner. Our analysis highlights novel challenges introduced by our benchmark and tasks, as well as numerous opportunities for future developments. We further broaden our analysis using a complementary real-world evaluation benchmark of in-situ navigation by pedestrians with disabilities. Our contributions provide an initial step towards widespread deployment of vision-based agents that can perceive and model the interaction needs of diverse people with disabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0db72788-0cc2-4a1f-a5b6-bc50dbf3bf19Cited by top-tier papers4
- XVO: Generalized Visual Odometry via Cross-Modal Self-TrainingLei Lai, Zhongkai Shangguan, Jimuyang Zhang, Eshed Ohn-BarICCV 2023 · 27 citations
- SelfD: Self-Learning Large-Scale Driving Policies From the WebJimuyang Zhang, Ruizhao Zhu, Eshed Ohn-BarCVPR 2022 · 17 citations
- Feedback-Guided Autonomous DrivingJimuyang Zhang, Zanming Huang, Arijit Ray, Eshed Ohn-BarCVPR 2024 · 15 citations
- Motion Diversification NetworksHee Jae Kim, Eshed Ohn-BarCVPR 2024
Builds on11
- Frustratingly Simple Few-Shot Object DetectionXin Wang, Thomas E. Huang, Joseph Gonzalez, Trevor Darrell et al.ICML 2020 · 723 citations
- Exploring the Limitations of Behavior Cloning for Autonomous DrivingFelipe Codevilla, Eder Santana, Antonio M. López, Adrien GaidonICCV 2019 · 666 citations
- Meta-Learning to Detect Rare ObjectsYu-Xiong Wang, Deva Ramanan, Martial HebertICCV 2019 · 339 citations
- nuScenes: A Multimodal Dataset for Autonomous DrivingHolger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora et al.CVPR 2020
- Sign Language Transformers: Joint End-to-End Sign Language Recognition and TranslationNecati Cihan Camgöz, Oscar Koller, Simon Hadfield, Richard BowdenCVPR 2020
Related papers
- Accessibility for Whom? Perceptions of Mobility Barriers Across Disability Groups and Implications for Designing Personalized MapsChu Li, Rock Yuren Pang, Delphine Labbé, Yochai Eisenberg et al.CHI 2025 · 15 citations
- R2H: Building Multimodal Navigation Helpers that Respond to Help RequestsYue Fan, Jing Gu, Kaizhi Zheng, Xin WangEMNLP 2023 · 3 citations
- Accessibility Scout: Personalized Accessibility Scans of Built EnvironmentsWilliam Huang, Xia Su, Jon E. Froehlich, Yang ZhangUIST 2025 · 3 citations
- Virtual Worlds Beyond Sight: Designing and Evaluating an Audio-Haptic System for Non-Visual VR ExplorationAayush Shrestha, Joseph MallochCHI 2025 · 6 citations
- From Selfie Stick to Virtual Cane: Enabling Blind Exploration through Mobile Virtual RealityHao Tang, Hong Zhao, Xinpeng Liu, Zhenchao Xia et al.CHI 2026
