Semantic Attention Flow Fields for Monocular Dynamic Scene Decomposition
Yiqing Liang, Eliot Laidlaw, Alexander Meyerowitz, Srinath Sridhar, James Tompkin
Abstract
From video, we reconstruct a neural volume that captures time-varying color, density, scene flow, semantics, and attention information. The semantics and attention let us identify salient foreground objects separately from the background across spacetime. To mitigate low resolution semantic and attention features, we compute pyramids that trade detail with whole-image context. After optimization, we perform a saliency-aware clustering to decompose the scene. To evaluate real-world scenes, we annotate object masks in the NVIDIA Dynamic Scene and DyCheck datasets. We demonstrate that this method can decompose dynamic scenes in an unsupervised way with competitive performance to a supervised method, and that it improves foreground/background segmentation over recent static/dynamic split methods. Project webpage: https://visual.cs.brown.edu/saff
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Grid4D: 4D Decomposed Hash Encoding for High-Fidelity Dynamic Gaussian SplattingJiawei Xu, Zexin Fan, Jian Yang, Jin XieNeurIPS 2024 · 64 citations
- When does perceptual alignment benefit vision representations?Shobhita Sundaram, Stephanie Fu, Lukas Muttenthaler, Netanel Tamir et al.NeurIPS 2024 · 24 citations
- WorDepth: Variational Language Prior for Monocular Depth EstimationZiyao Zeng, Daniel Wang, Fengyu Yang, Hyoungseob Park et al.CVPR 2024 · 20 citations
- DASH: 4D Hash Encoding with Self-Supervised Decomposition for Real-Time Dynamic Scene RenderingJie Chen, Zhangchi Hu, Peixi Wu, Huyue Zhu et al.ICCV 2025 · 2 citations
- DIV-FF: Dynamic Image-Video Feature Fields For Environment Understanding in Egocentric VideosLorenzo Mur-Labadia, Josechu Guerrero, Ruben Martinez-CantinCVPR 2025
Builds on28
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman et al.ICCV 2021 · 2,700 citations
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- Non-Rigid Neural Radiance Fields: Reconstruction and Novel View Synthesis of a Dynamic Scene From Monocular VideoEdgar Tretschk, Ayush Tewari, Vladislav Golyanik, Michael Zollhöfer et al.ICCV 2021 · 617 citations
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
Related papers
- Semantic Flow: Learning Semantic Fields of Dynamic Scenes from Monocular VideosFengrui Tian, Yueqi Duan, Angtian Wang, Jianfei Guo et al.ICLR 2024 · 7 citations
- DynaVol: Unsupervised Learning for Dynamic Scenes through Object-Centric VoxelizationYanpeng Zhao, Siyu Gao, Yunbo Wang, Xiaokang YangICLR 2024 · 2 citations
- DS-NeRV: Implicit Neural Video Representation with Decomposed Static and Dynamic CodesHao Yan, Zhihui Ke, Xiaobo Zhou, Tie Qiu et al.CVPR 2024 · 18 citations
- DeGauss: Dynamic-Static Decomposition with Gaussian Splatting for Distractor-Free 3D ReconstructionRui Wang, Quentin Lohmeyer, Mirko Meboldt, Siyu TangICCV 2025 · 12 citations
- NVFi: Neural Velocity Fields for 3D Physics Learning from Dynamic VideosJinxi Li, Ziyang Song, Bo YangNeurIPS 2023 · 39 citations
