MAAL: Multimodality-Aware Autoencoder-based Affordance Learning for 3D Articulated Objects
Yuanzhi Liang, Xiaohan Wang, Linchao Zhu, Yi Yang
摘要
Inferring affordance for 3D articulated objects is a challenging and practical problem. It is a primary problem for applying robots to real-world scenarios. The exploration can be summarized as figuring out where to act and how to act. Correspondingly, the task mainly requires producing actionability scores, action proposals, and success likelihood scores according to the given 3D object information and robotic information. Current works usually directly process multi-modal inputs with early fusion and apply critic networks to produce scores, which leads to insufficient multi-modal learning ability and inefficiently iterative training in multiple stages. This paper proposes a novel Multimodality-Aware Autoencoder-based affordance Learning (MAAL) for the 3D object affordance problem. It is an efficient pipeline, trained in one go, and only requires a few positive samples in training data. More importantly, MAAL contains a MultiModal Energized Encoder (MME) for better multi-modal learning. It comprehensively models all multi-modal inputs from 3D objects and robotic actions. Jointly considering information from multiple modalities, the encoder further learns interactions between robots and objects. MME empowers the better multi-modal learning ability for understanding object affordance. Experimental results and visualizations, based on a large-scale dataset PartNet-Mobility, show the effectiveness of MAAL in learning multi-modal data and solving the 3D articulated object affordance problem.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- An Interactive Navigation Method with Effect-oriented AffordanceXiaohan Wang, Yuehu Liu, Xinhang Song, Yuyi Liu 等CVPR 2024 · 被引用 1 次
- Vision-Guided Action: Enhancing 3D Human Motion Prediction with Gaze-informed Affordance in 3D ScenesTing Yu, Yi Lin, Jun Yu, Zhenyu Lou 等CVPR 2025
- Weakly-Supervised Affordance Grounding Guided by Part-Level Semantic PriorsPeiran Xu, Yadong MuICLR 2025
它引用的顶会 Paper10
- Hand-Object Contact Consistency Reasoning for Human Grasps GenerationHanwen Jiang, Shaowei Liu, Jiashun Wang, Xiaolong WangICCV 2021 · 被引用 242 次
- Where2Act: From Pixels to Actions for Articulated 3D ObjectsKaichun Mo, Leonidas J. Guibas, Mustafa Mukadam, Abhinav Gupta 等ICCV 2021 · 被引用 240 次
- OmniVL: One Foundation Model for Image-Language and Video-Language TasksJunke Wang, Dongdong Chen, Zuxuan Wu, Chong Luo 等NeurIPS 2022 · 被引用 205 次
- VrR-VG: Refocusing Visually-Relevant RelationshipsYuanzhi Liang, Yalong Bai, Wei Zhang, Xueming Qian 等ICCV 2019 · 被引用 93 次
- SEEG: Semantic Energized Co-speech Gesture GenerationYuanzhi Liang, Qianyu Feng, Linchao Zhu, Li Hu 等CVPR 2022 · 被引用 53 次
相关 Paper
- DualAfford: Learning Collaborative Visual Affordance for Dual-gripper ManipulationYan Zhao, Ruihai Wu, Zhehuan Chen, Yourong Zhang 等ICLR 2023 · 被引用 2 次
- Grounding 3D Object Affordance with Language Instructions, Visual Observations and InteractionsHe Zhu, Quyu Kong, Kechun Xu, Xunlong Xia 等CVPR 2025
- VAT-Mart: Learning Visual Action Trajectory Proposals for Manipulating 3D ARTiculated ObjectsRuihai Wu, Yan Zhao, Kaichun Mo, Zizheng Guo 等ICLR 2022 · 被引用 119 次
- 3D AffordanceNet: A Benchmark for Visual Object Affordance UnderstandingShengheng Deng, Xun Xu, Chaozheng Wu, Ke Chen 等CVPR 2021
- Interactive Anomaly Detection for Articulated Objects via Motion AnticipationAnkan Bhunia, Changjian Li, Hakan BilenNeurIPS 2025 · 被引用 1 次
