STORM: Segment, Track, and Object Re-Localization from a Single Image
Yu Deng, Teng Cao, Hikaru Shindo, Quentin Delfosse, Jiahong Xue, Kristian Kersting
Abstract
Accurate 6D pose estimation and tracking are core capabilities for physical AI systems, yet real-world deployment remains brittle and labor-intensive. Many pipelines rely on CAD models, manual masking, or per-object adaptation, and still fail under occlusion or fast motion without a principled way to recognize failure. We propose STORM, a unified framework for reference-conditioned 6D tracking that can operate from a single reference image, with minimal manual input and improved robustness. STORM combines: (i) Hierarchical Spatial Fusion Attention (HSFA), a task-driven reference-query fusion architecture that supports both single-reference and multi-reference conditioning and can optionally use vision-language semantic conditioning to resolve instance ambiguities; and (ii) a BCE-trained tracking verifier whose continuous compatibility logit is used as an energy-like score to detect drift and trigger automatic re-initialization. Experiments on LM-O and YCB-Video show that STORM improves annotation-free pose tracking accuracy over strong baselines and recovers reliably from severe occlusions and rapid viewpoint changes with minimal overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 232e8ee7-2dd3-451e-a9cc-f269612e5ea7Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 2,213 citations
- Personalize Segment Anything Model with One ShotRenrui Zhang, Zhengkai Jiang, Ziyu Guo, Shilin Yan et al.ICLR 2024 · 333 citations
- FoundationPose: Unified 6D Pose Estimation and Tracking of Novel ObjectsBowen Wen, Wei Yang, Jan Kautz, Stan BirchfieldCVPR 2024 · 215 citations
- ZebraPose: Coarse to Fine Surface Encoding for 6DoF Object Pose EstimationYongzhi Su, Mahdi Saleh, Torben Fetzer, Jason R. Rambach et al.CVPR 2022 · 170 citations
Related papers
- One2Any: One-Reference 6D Pose Estimation for Any ObjectMengya Liu, Siyuan Li, Ajad Chhatkuli, Prune Truong et al.CVPR 2025
- Optical Flow-Guided 6DoF Object Pose Tracking with an Event CameraZibin Liu, Banglei Guan, Yang Shang, Shunkun Liang et al.ACM MM 2024 · 4 citations
- KV-Tracker: Real-Time Pose Tracking with TransformersMarwan Taher, Ignacio Alzugaray, Kirill Mazur, Xin Kong et al.CVPR 2026 · 5 citations
- Physical Inertial Poser (PIP): Physics-aware Real-time Human Motion Tracking from Sparse Inertial SensorsXinyu Yi, Yuxiao Zhou, Marc Habermann, Soshi Shimada et al.CVPR 2022 · 198 citations
- What Happens Next? Anticipating Future Motion by Generating Point TrajectoriesGabrijel Boduljak, Laurynas Karazija, Iro Laina, Christian Rupprecht et al.ICLR 2026 · 10 citations
