E-MaT: Event-oriented Mamba for Egocentric Point Tracking
Han Han, Wei Zhai, Baocai Yin, Yang Cao, Bin Li, Zhengjun Zha
摘要
Egocentric point tracking aims to localize points on object surfaces from a first-person perspective and serves as a critical step toward embodied intelligence. Recent methods rely on video input, tracking query points through feature matching across consecutive frames. However, these methods struggle in highly dynamic settings—a common challenge in first-person perspectives, where the head-mounted camera undergoes frequent and abrupt rotations, resulting in high angular velocities, motion blur, and large inter-frame displacements. In contrast, event cameras capture motion at microsecond temporal resolution, naturally avoiding blur and delivering low-latency, high-fidelity cues crucial for egocentric point tracking. Moreover, rapid egocentric motion disrupts local smoothness, breaking the assumption that spatially adjacent regions share similar motion. Event dynamics expose global motion trends, guiding coherent modeling and consistent feature flow. Therefore, this paper proposes a mamba-based tracking framework that constructs feature modeling paths aligned with the dominant motion trend extracted from events, and modulates feature propagation along these paths based on local motion intensity, enhancing stability by suppressing unreliable signals and emphasizing consistent cues. Additionally, a motion-adaptive suppression module enhances temporal robustness by adaptively suppressing correlation features based on motion intensity variations, mitigating the effects of intensity fluctuations and partial observability. To facilitate research in this domain, a multimodal dataset named DVS-EgoPoints with both events and videos for egocentric point tracking is collected. Experiments on the DVS-EgoPoints dataset and a simulation benchmark demonstrate superior performance over state-of-the-art methods, especially under challenging motion and occlusion conditions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang 等ICML 2024 · 被引用 1,725 次
- PointOdyssey: A Large-Scale Synthetic Dataset for Long-Term Point TrackingYang Zheng, Adam W. Harley, Bokui Shen, Gordon Wetzstein 等ICCV 2023 · 被引用 255 次
- Multi-grained Spatio-Temporal Features Perceived Network for Event-based Lip-ReadingGanchao Tan, Yang Wang, Han Han, Yang Cao 等CVPR 2022 · 被引用 36 次
- EF-3DGS: Event-Aided Free-Trajectory 3D Gaussian SplattingBohao Liao, Wei Zhai, Zengyu Wan, Zhixin Cheng 等NeurIPS 2025 · 被引用 19 次
- TAPTRv3: Spatial and Temporal Context Foster Robust Tracking of Any Point in Long VideoJinyuan Qu, Hongyang Li, Shilong Liu, Tianhe Ren 等ICLR 2026 · 被引用 9 次
相关 Paper
- MATE: Motion-Augmented Temporal Consistency for Event-Based Point TrackingHan Han, Wei Zhai, Yang Cao, Bin Li 等ICCV 2025 · 被引用 3 次
- ETAP: Event-based Tracking of Any PointFriedhelm Hamann, Daniel Gehrig, Filbert Febryanto, Kostas Daniilidis 等CVPR 2025
- EventEgo3D: 3D Human Motion Capture from Egocentric Event StreamsChristen Millerdurai, Hiroyasu Akada, Jian Wang, Diogo C. Luvizon 等CVPR 2024 · 被引用 11 次
- E-3DPSM: A State Machine for Event-based Egocentric 3D Human Pose EstimationMayur Deshmukh, Hiroyasu Akada, Helge Rhodin, Christian Theobalt 等CVPR 2026 · 被引用 1 次
- Event6D: Event-based Novel Object 6D Pose TrackingJae-Young Kang, Hoonhee Cho, Taeyeop Lee, Minjun Kang 等CVPR 2026 · 被引用 4 次
