What If Smart Homes Could See Our Homes?: Exploring DIY Smart Home Building Experiences with VLM-Based Camera Sensors
Sojeong Yun, Youn-kyung Lim
Abstract
The advancement of Vision-Language Model (VLM) camera sensors, which enable autonomous understanding of household situations without user intervention, has the potential to completely transform the DIY smart home building experience. Will this simplify or complicate the DIY smart home process? Additionally, what features do users want to create using these sensors? To explore this, we conducted a three-week diary-based experience prototyping study with 12 participants. Participants recorded their daily activities, used GPT to analyze the images, and manually customized and tested smart home features based on the analysis. The study revealed three key findings: (1) participants' expectations for VLM camera-based smart homes, (2) the impact of VLM camera sensor characteristics on the DIY process, and (3) users' concerns. Through the findings of this study, we propose design implications to support the DIY smart home building process with VLM camera sensors, and discuss living with intelligence.
• Human-centered computing → Empirical studies in ubiquitous and mobile computing; Empirical studies in interaction design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fb96d3e5-305d-48d5-9453-c24f24caab98Cited by top-tier papers3
- Balanced Token Pruning: Accelerating Vision Language Models Beyond Local OptimizationKaiyuan Li, Xiaoyue Chen, Chen Gao, Yong Li et al.NeurIPS 2025 · 26 citations
- Situated, Dynamic, and Subjective: Envisioning the Design of Theory-of-Mind-Enabled Everyday AI with Industry PractitionersQiaosi Wang, Jini Kim, Avanita Sharma, Alicia (Hyun Jin) Lee et al.CHI 2026 · 1 citation
- OrganicHAR: Towards Activity Discovery in Organic Settings for Privacy Preserving Sensors Using Efficient Video AnalysisPrasoon Patidar, Riku Arakawa, Ricardo Graça, Ruben Moutinho et al.UbiComp 2026
Builds on32
- "It's Kind of Like Code-Switching": Black Older Adults' Experiences with a Voice Assistant for Health Information SeekingChristina N. Harrington, Radhika Garg, Amanda T. Woodward, Dimitri WilliamsCHI 2022 · 99 citations
- Sasha: Creative Goal-Oriented Reasoning in Smart Homes with Large Language ModelsEvan King, Haoxiang Yu, Sangsu Lee, Christine JulienUbiComp 2024 · 91 citations
- GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented RealityJaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu et al.CHI 2024 · 86 citations
- Multimodal Healthcare AI: Identifying and Designing Clinically Relevant Vision-Language Applications for RadiologyNur Yildirim, Hannah Richardson, Maria Teodora Wetscherek, Junaid Bajwa et al.CHI 2024 · 81 citations
- "It did not give me an option to decline": A Longitudinal Analysis of the User Experience of Security and Privacy in Smart Home ProductsGeorge Chalhoub, Martin J. Kraemer, Norbert Nthala, Ivan FlechaisCHI 2021 · 71 citations
Related papers
- Potential and Challenges of DIY Smart Homes with an ML-intensive Camera SensorSojeong Yun, Youn-Kyung LimCHI 2023 · 9 citations
- Peeking Ahead of the Field Study: Exploring VLM Personas as Support Tools for Embodied Studies in HCIXinyue Gui, Ding Xia, Mark Colley, Yuan Li et al.CHI 2026 · 1 citation
- Not Seeing the Whole Picture: Challenges and Opportunities in Using AI for Co-Making Physical, DIY-AT for People with Visual ImpairmentsBen Kosa, Hsuanling Lee, Jasmine Li, Sanbrita Mondal et al.CHI 2026 · 1 citation
- Behavior-Aware Anthropometric Scene Generation for Human-Usable 3D LayoutsSemin Jin, Donghyuk Kim, Jeongmin Ryu, Kyung Hoon HyunCHI 2026 · 1 citation
- ImagineNav: Prompting Vision-Language Models as Embodied Navigator through Scene ImaginationXinxin Zhao, Wenzhe Cai, Likun Tang, Teng WangICLR 2025
