DIV-FF: Dynamic Image-Video Feature Fields For Environment Understanding in Egocentric Videos
Lorenzo Mur-Labadia, Josechu Guerrero, Ruben Martinez-Cantin
2025Year
Abstract
Figure 1. DIV-FF distills image and video language features in a triple stream feature field tailored to egocentric videos with numerous interactions and camera wearer movements. Our approach achieves a deep understanding of the environment, supporting precise affordance segmentation, semantic scene decomposition and consistent segmentation of dynamic objects. With its implicit 3D representation, DIV-FF comprehends not just novel views but also surrounding areas.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- LERF: Language Embedded Radiance FieldsJustin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa et al.ICCV 2023 · 620 citations
Related papers
- EA3D: Online Open-World 3D Object Extraction from Streaming VideosXiaoyu Zhou, Jingqi Wang, Yuang Jia, Yongtao Wang et al.NeurIPS 2025 · 5 citations
- VLScene: Vision-Language Guidance Distillation for Camera-Based 3D Semantic Scene CompletionMeng Wang, Huilong Pi, Ruihui Li, Yunchuan Qin et al.AAAI 2025 · 11 citations
- DeGauss: Dynamic-Static Decomposition with Gaussian Splatting for Distractor-Free 3D ReconstructionRui Wang, Quentin Lohmeyer, Mirko Meboldt, Siyu TangICCV 2025 · 12 citations
- DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular VideosWen-Hsuan Chu, Lei Ke, Katerina FragkiadakiNeurIPS 2024 · 75 citations
- Decoupling Static and Hierarchical Motion Perception for Referring Video SegmentationShuting He, Henghui DingCVPR 2024
