What If Smart Homes Could See Our Homes?: Exploring DIY Smart Home Building Experiences with VLM-Based Camera Sensors
Sojeong Yun, Youn-kyung Lim
摘要
The advancement of Vision-Language Model (VLM) camera sensors, which enable autonomous understanding of household situations without user intervention, has the potential to completely transform the DIY smart home building experience. Will this simplify or complicate the DIY smart home process? Additionally, what features do users want to create using these sensors? To explore this, we conducted a three-week diary-based experience prototyping study with 12 participants. Participants recorded their daily activities, used GPT to analyze the images, and manually customized and tested smart home features based on the analysis. The study revealed three key findings: (1) participants' expectations for VLM camera-based smart homes, (2) the impact of VLM camera sensor characteristics on the DIY process, and (3) users' concerns. Through the findings of this study, we propose design implications to support the DIY smart home building process with VLM camera sensors, and discuss living with intelligence.
• Human-centered computing → Empirical studies in ubiquitous and mobile computing; Empirical studies in interaction design.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Balanced Token Pruning: Accelerating Vision Language Models Beyond Local OptimizationKaiyuan Li, Xiaoyue Chen, Chen Gao, Yong Li 等NeurIPS 2025 · 被引用 26 次
- Situated, Dynamic, and Subjective: Envisioning the Design of Theory-of-Mind-Enabled Everyday AI with Industry PractitionersQiaosi Wang, Jini Kim, Avanita Sharma, Alicia (Hyun Jin) Lee 等CHI 2026 · 被引用 1 次
- OrganicHAR: Towards Activity Discovery in Organic Settings for Privacy Preserving Sensors Using Efficient Video AnalysisPrasoon Patidar, Riku Arakawa, Ricardo Graça, Ruben Moutinho 等UbiComp 2026
它引用的顶会 Paper32
- "It's Kind of Like Code-Switching": Black Older Adults' Experiences with a Voice Assistant for Health Information SeekingChristina N. Harrington, Radhika Garg, Amanda T. Woodward, Dimitri WilliamsCHI 2022 · 被引用 99 次
- Sasha: Creative Goal-Oriented Reasoning in Smart Homes with Large Language ModelsEvan King, Haoxiang Yu, Sangsu Lee, Christine JulienUbiComp 2024 · 被引用 91 次
- GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented RealityJaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu 等CHI 2024 · 被引用 86 次
- Multimodal Healthcare AI: Identifying and Designing Clinically Relevant Vision-Language Applications for RadiologyNur Yildirim, Hannah Richardson, Maria Teodora Wetscherek, Junaid Bajwa 等CHI 2024 · 被引用 81 次
- "It did not give me an option to decline": A Longitudinal Analysis of the User Experience of Security and Privacy in Smart Home ProductsGeorge Chalhoub, Martin J. Kraemer, Norbert Nthala, Ivan FlechaisCHI 2021 · 被引用 71 次
相关 Paper
- Potential and Challenges of DIY Smart Homes with an ML-intensive Camera SensorSojeong Yun, Youn-Kyung LimCHI 2023 · 被引用 9 次
- Peeking Ahead of the Field Study: Exploring VLM Personas as Support Tools for Embodied Studies in HCIXinyue Gui, Ding Xia, Mark Colley, Yuan Li 等CHI 2026 · 被引用 1 次
- Not Seeing the Whole Picture: Challenges and Opportunities in Using AI for Co-Making Physical, DIY-AT for People with Visual ImpairmentsBen Kosa, Hsuanling Lee, Jasmine Li, Sanbrita Mondal 等CHI 2026 · 被引用 1 次
- Behavior-Aware Anthropometric Scene Generation for Human-Usable 3D LayoutsSemin Jin, Donghyuk Kim, Jeongmin Ryu, Kyung Hoon HyunCHI 2026 · 被引用 1 次
- ImagineNav: Prompting Vision-Language Models as Embodied Navigator through Scene ImaginationXinxin Zhao, Wenzhe Cai, Likun Tang, Teng WangICLR 2025
