RealAppiance: Let High-fidelity Appliance Assets Controllable and Workable as Aligned Real Manauls
Yuzheng Gao, Yuxing Long, Lei Kang, Yuchong Guo, Ziyan Yu, Shangqing Mao, Jiyao Zhang, Ruihai Wu, Dongjiang Li, Hui Shen, Hao Dong
Abstract
Existing appliance assets suffer from poor rendering, incomplete mechanisms, and misalignment with manuals, leading to simulation-reality gaps that hinder appliance manipulation development. In this work, we introduce the RealAppliance dataset, comprising 100 high-fidelity appliances with complete physical, electronic mechanisms, and program logic aligned with their manuals. Based on these assets, we propose the RealAppliance-Bench benchmark, which evaluates multimodal large language models and embodied manipulation planning models across key tasks in appliance manipulation planning: manual page retrieval, appliance part grounding, open-loop manipulation planning, and closed-loop planning adjustment. Our analysis of model performances on RealAppliance-Bench provides insights for advancing appliance manipulation research
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- ArtVIP: Articulated Digital Assets of Visual Realism, Modular Interaction, and Physical Fidelity for Robot LearningZhao Jin, Zhengping Che, Tao Li, Zhen Zhao et al.ICLR 2026 · 15 citations
- SAPIEN: A SimulAted Part-Based Interactive ENvironmentFanbo Xiang, Yuzhe Qin, Kaichun Mo, Yikuan Xia et al.CVPR 2020
- ManipLLM: Embodied Multimodal Large Language Model for Object-Centric Robotic ManipulationXiaoqi Li, Mingxu Zhang, Yiran Geng, Haoran Geng et al.CVPR 2024
- CheckManual: A New Challenge and Benchmark for Manual-based Appliance ManipulationYuxing Long, Jiyao Zhang, Mingjie Pan, Tianshu Wu et al.CVPR 2025
Related papers
- Placeit3d: Language-Guided Object Placement in Real 3D ScenesAhmed Abdelreheem, Filippo Aleotti, Jamie Watson, Zawar Qureshi et al.ICCV 2025 · 11 citations
- AmbiK: Dataset of Ambiguous Tasks in Kitchen EnvironmentAnastasiia Ivanova, Eva Bakaeva, Zoya Volovikova, Alexey K. Kovalev et al.ACL 2025
- AIR-VLA: Vision-Language-Action Systems for Aerial ManipulationJianli Sun, Bin Tian, Qiyao Zhang, Chengxiang Li et al.ICML 2026 · 4 citations
- Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal AgentsTianyi Men, Zhuoran Jin, Pengfei Cao, Yubo Chen et al.ACL 2025 · 13 citations
- PAI-Bench: A Comprehensive Benchmark For Physical AIFengzhe Zhou, Jiannan Huang, Jialuo Li, Deva Ramanan et al.CVPR 2026 · 32 citations
