DIV-FF: Dynamic Image-Video Feature Fields For Environment Understanding in Egocentric Videos
Lorenzo Mur-Labadia, Josechu Guerrero, Ruben Martinez-Cantin
2025年份
摘要
Figure 1. DIV-FF distills image and video language features in a triple stream feature field tailored to egocentric videos with numerous interactions and camera wearer movements. Our approach achieves a deep understanding of the environment, supporting precise affordance segmentation, semantic scene decomposition and consistent segmentation of dynamic objects. With its implicit 3D representation, DIV-FF comprehends not just novel views but also surrounding areas.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- LERF: Language Embedded Radiance FieldsJustin Kerr, Chung Min Kim, Ken Goldberg, Angjoo Kanazawa 等ICCV 2023 · 被引用 620 次
相关 Paper
- EA3D: Online Open-World 3D Object Extraction from Streaming VideosXiaoyu Zhou, Jingqi Wang, Yuang Jia, Yongtao Wang 等NeurIPS 2025 · 被引用 5 次
- VLScene: Vision-Language Guidance Distillation for Camera-Based 3D Semantic Scene CompletionMeng Wang, Huilong Pi, Ruihui Li, Yunchuan Qin 等AAAI 2025 · 被引用 11 次
- DeGauss: Dynamic-Static Decomposition with Gaussian Splatting for Distractor-Free 3D ReconstructionRui Wang, Quentin Lohmeyer, Mirko Meboldt, Siyu TangICCV 2025 · 被引用 12 次
- DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular VideosWen-Hsuan Chu, Lei Ke, Katerina FragkiadakiNeurIPS 2024 · 被引用 75 次
- Decoupling Static and Hierarchical Motion Perception for Referring Video SegmentationShuting He, Henghui DingCVPR 2024
