Multi-Object Manipulation via Object-Centric Neural Scattering Functions
Stephen Tian, Yancheng Cai, Hong-Xing Yu, Sergey Zakharov, Katherine Liu, Adrien Gaidon, Yunzhu Li, Jiajun Wu
Abstract
Learned visual dynamics models have proven effective for robotic manipulation tasks. Yet, it remains unclear how best to represent scenes involving multi-object interactions. Current methods decompose a scene into discrete objects, but they struggle with precise modeling and manipulation amid challenging lighting conditions as they only encode appearance tied with specific illuminations. In this work, we propose using object-centric neural scattering functions (OSFs) as object representations in a model-predictive control framework. OSFs model per-object light transport, enabling compositional scene re-rendering under object rearrangement and varying lighting conditions. By combining this approach with inverse parameter estimation and graph-based neural dynamics models, we demonstrate improved model-predictive control performance and generalization in compositional multi-object environments, even in previously unseen scenarios and harsh lighting conditions. * indicates equal contribution. Yancheng is affiliated with Fudan University; this work was done while he was a summer intern at Stanford.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Inferring Hybrid Neural Fluid Fields from VideosHong-Xing Yu, Yang Zheng, Yuan Gao, Yitong Deng et al.NeurIPS 2023 · 36 citations
- DEL: Discrete Element Learner for Learning 3D Particle Dynamics with Neural RenderingJiaxu Wang, Jingkai Sun, Ziyi Zhang, Junhao He et al.NeurIPS 2024 · 5 citations
- Digital Twin Catalog: A Large-Scale Photorealistic 3D Object Digital Twin DatasetZhao Dong, Ka Chen, Zhaoyang Lv, Hong-Xing Yu et al.CVPR 2025
- Do Computer Vision Foundation Models Learn the Low-level Characteristics of the Human Visual System?Yancheng Cai, Fei Yin, Dounia Hammou, Rafal MantiukCVPR 2025
Builds on14
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- KiloNeRF: Speeding up Neural Radiance Fields with Thousands of Tiny MLPsChristian Reiser, Songyou Peng, Yiyi Liao, Andreas GeigerICCV 2021 · 963 citations
- Pix2Pose: Pixel-Wise Coordinate Regression of Objects for 6D Pose EstimationKiru Park, Timothy Patten, Markus VinczeICCV 2019 · 527 citations
- DPOD: 6D Pose Object Detector and RefinerSergey Zakharov, Ivan Shugurov, Slobodan IlicICCV 2019 · 486 citations
Related papers
- Object-Centric Representation Learning with Generative Spatial-Temporal FactorizationNanbo Li, Muhammad Ahmed Raza, Wenbin Hu, Zhaole Sun et al.NeurIPS 2021 · 17 citations
- DynaVol: Unsupervised Learning for Dynamic Scenes through Object-Centric VoxelizationYanpeng Zhao, Siyu Gao, Yunbo Wang, Xiaokang YangICLR 2024 · 2 citations
- LightFormer: Light-Oriented Global Neural Rendering in Dynamic SceneHaocheng Ren, Yuchi Huo, Yifan Peng, Hongtao Sheng et al.SIGGRAPH 2024 · 7 citations
- Neural Scene Graphs for Dynamic ScenesJulian Ost, Fahim Mannan, Nils Thuerey, Julian Knodt et al.CVPR 2021
- DM-NeRF: 3D Scene Geometry Decomposition and Manipulation from 2D ImagesBing Wang, Lu Chen, Bo YangICLR 2023 · 26 citations
