Disability-First Design and Creation of A Dataset Showing Private Visual Information Collected With People Who Are Blind
Tanusree Sharma, Abigale Stangl, Lotus Zhang, Yu-Yun Tseng, Inan Xu, Leah Findlater, Danna Gurari, Yang Wang
Abstract
We present the design and creation of a disability-first dataset, “BIV-Priv,” which contains 728 images and 728 videos of 14 private categories captured by 26 blind participants to support downstream development of artificial intelligence (AI) models. While best practices in dataset creation typically attempt to eliminate private content, some applications require such content for model development. We describe our approach in creating this dataset with private content in an ethical way, including using props rather than participants’ own private objects and balancing multi-disciplinary perspectives (e.g., accessibility, privacy, computer vision) to meet the tangible metrics (e.g., diversity, category, amount of content) to support AI innovations. We observed challenges that our participants encountered during the data collection, including accessibility issues (e.g., understanding foreground vs. background object placement) and issues due to the sensitive nature of the content (e.g., discomfort in capturing some props such as condoms around family members).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 51faf0df-08cb-4b62-9603-6b8d1f2ef317Cited by top-tier papers9
- Everyday Uncertainty: How Blind People Use GenAI Tools for Information AccessXinru Tang, Ali Abdolrahmani, Darren Gergle, Anne Marie PiperCHI 2025 · 26 citations
- Examining Human Perception of Generative Content Replacement in Image Privacy ProtectionAnran Xu, Shitao Fang, Huan Yang, Simo Hosio et al.CHI 2024 · 22 citations
- DIPA2: An Image Dataset with Cross-cultural Privacy Perception AnnotationsAnran Xu, Zhongyi Zhou, Kakeru Miyazaki, Ryo Yoshikawa et al.UbiComp 2024 · 15 citations
- Exploring the Experiences of Individuals Who are Blind or Low-Vision Using Object-Recognition Technologies in IndiaGesu India, Simon Robinson, Jennifer Pearson, Cecily Morrison et al.CHI 2025 · 8 citations
- Disability-First AI Dataset Annotation: Co-designing Stuttered Speech Annotation Guidelines with People Who StutterXinru Tang, Jingjin Li, Shaomei WuCHI 2026 · 3 citations
Builds on13
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong et al.CHI 2021 · 725 citations
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- Feature Weighting and Boosting for Few-Shot SegmentationKhoi Nguyen, Sinisa TodorovicICCV 2019 · 402 citations
- How We've Taught Algorithms to See Identity: Constructing Race and Gender in Image Databases for Facial AnalysisMorgan Klaus Scheuerman, Kandrea Wade, Caitlin Lustig, Jed R. BrubakerCSCW 2020 · 198 citations
- ORBIT: A Real-World Few-Shot Dataset for Teachable Object RecognitionDaniela Massiceti, Luisa M. Zintgraf, John Bronskill, Lida Theodorou et al.ICCV 2021 · 55 citations
Related papers
- Contributing to Accessibility Datasets: Reflections on Sharing Study Data by Blind PeopleRie Kamikubo, Kyungjun Lee, Hernisa KacorriCHI 2023 · 16 citations
- VideoA11y: Method and Dataset for Accessible Video DescriptionChaoyu Li, Sid Padmanabhuni, Maryam S. Cheema, Hasti Seifi et al.CHI 2025 · 23 citations
- VisAssist: A Visually Impaired-Captured Video Question Answering Benchmark for Assistive SystemsQi Gao, Heng Li, Yixin Zhou, Meixuan Zhou et al.AAAI 2026
- The ORBIT India Dataset: Understanding the Challenges of Collecting a Disability-First AI Dataset in Low-Resource EnvironmentsGesu India, Martin Grayson, Cecily Morrison, Daniela Massiceti et al.CHI 2026 · 1 citation
- A New Dataset Based on Images Taken by Blind People for Testing the Robustness of Image Classification Models Trained for ImageNet CategoriesReza Akbarian Bafghi, Danna GurariCVPR 2023
