Multi-Object Discovery by Low-Dimensional Object Motion
Sadra Safadoust, Fatma Güney
Abstract
Recent work in unsupervised multi-object segmentation shows impressive results by predicting motion from a single image despite the inherent ambiguity in predicting motion without the next image. On the other hand, the set of possible motions for an image can be constrained to a low-dimensional space by considering the scene structure and moving objects in it. We propose to model pixel-wise geometry and object motion to remove ambiguity in reconstructing flow from a single image. Specifically, we divide the image into coherently moving regions and use depth to construct flow bases that best explain the observed flow in each region. We achieve state-of-the-art results in unsupervised multi-object segmentation on synthetic and real-world datasets by modeling the scene structure and object motion. Our evaluation of the predicted depth maps shows reliable performance in monocular depth estimation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 62612ccd-40a1-43bf-8170-1b79b19bb70eCited by top-tier papers9
- Object-Centric Learning for Real-World Videos by Predicting Temporal Feature SimilaritiesAndrii Zadaianchuk, Maximilian Seitzer, Georg MartiusNeurIPS 2023 · 104 citations
- Learning Segmentation from Point TrajectoriesLaurynas Karazija, Iro Laina, Christian Rupprecht, Andrea VedaldiNeurIPS 2024 · 14 citations
- FlowSeek: Optical Flow Made Easier with Depth Foundation Models and Motion BasesMatteo Poggi, Fabio TosiICCV 2025 · 5 citations
- Predicting Video Slot Attention Queries from Random Slot-Feature PairsRongzhen Zhao, Jian Li, Juho Kannala, Joni PajarinenAAAI 2026 · 3 citations
- Scene-Centric Unsupervised Video Panoptic SegmentationChristoph Reich, Oliver Hahn, Nikita Araslanov, Laura Leal-Taixe et al.CVPR 2026 · 1 citation
Builds on24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- GENESIS: Generative Scene Inference and Sampling with Object-Centric Latent RepresentationsMartin Engelcke, Adam R. Kosiorek, Oiwi Parker Jones, Ingmar PosnerICLR 2020 · 334 citations
Related papers
- AdaDepth: Exploiting Inherent Scene Information for Self-Supervised Depth Estimation in Dynamic ScenesXuanang Gao, Xiongbin Wu, Zhiwei Ning, Runze Yang et al.AAAI 2026
- Dynamo-Depth: Fixing Unsupervised Depth Estimation for Dynamical ScenesYihong Sun, Bharath HariharanNeurIPS 2023 · 58 citations
- Mono-SF: Multi-View Geometry Meets Single-View Depth for Monocular Scene Flow Estimation of Dynamic Traffic ScenesFabian Brickwedde, Steffen Abraham, Rudolf MesterICCV 2019 · 55 citations
- SynDeMo: Synergistic Deep Feature Alignment for Joint Learning of Depth and Ego-MotionBehzad Bozorgtabar, Mohammad Saeed Rad, Dwarikanath Mahapatra, Jean-Philippe ThiranICCV 2019 · 44 citations
- RM-Depth: Unsupervised Learning of Recurrent Monocular Depth in Dynamic ScenesTak-Wai HuiCVPR 2022 · 62 citations
