RMem: Restricted Memory Banks Improve Video Object Segmentation
Junbao Zhou, Ziqi Pang, Yu-Xiong Wang
摘要
With recent video object segmentation (VOS) benchmarks evolving to challenging scenarios, we revisit a sim-ple but overlooked strategy: restricting the size of memory banks. This diverges from the prevalent practice of ex-panding memory banks to accommodate extensive histor-ical information. Our specially designed “memory deci-phering” study offers a pivotal insight underpinning such a strategy: expanding memory banks, while seemingly bene-ficial, actually increases the difficulty for VOS modules to decode relevant features due to the confusion from redun-dant information. By restricting memory banks to a limited number of essential frames, we achieve a notable improvement in VOS accuracy. This process balances the im-portance and freshness of frames to maintain an informative memory bank within a bounded capacity. Additionally, restricted memory banks reduce the training-inference discrepancy in memory lengths compared with continuous expansion. This fosters new opportunities in temporal reasoning and enables us to introduce the previously overlooked “temporal positional embedding.” Finally, our insights are embodied in “RMem” (“R” for restricted), a simple yet effective VOS modification that excels at challenging VOS scenarios and establishes new state of the art for object state changes (on the VOST dataset) and long videos (on the Long Videos dataset). Our code and demos are available at https://restricted-memory.github.io/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic ManipulationHao Shi, Bin Xie, Yingfei Liu, Lin Sun 等ICLR 2026 · 被引用 227 次
- Advancing Complex Video Object Segmentation via Progressive Concept ConstructionZhixiong Zhang, Shuangrui Ding, Xiaoyi Dong, Songxin He 等ICLR 2026 · 被引用 17 次
- SAM2LONG: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory TreeShuangrui Ding, Rui Qian, Xiaoyi Dong, Pan Zhang 等ICCV 2025 · 被引用 15 次
- Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!Junbao Zhou, Yuan Zhou, Kesen Zhao, Qingshan Xu 等ICLR 2026 · 被引用 7 次
- One Token per Highly Selective Frame: Towards Extreme Compression for Long Video UnderstandingZheyu Zhang, Ziqi Pang, Shixing Chen, Xiang Hao 等NeurIPS 2025 · 被引用 5 次
它引用的顶会 Paper21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 被引用 845 次
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object SegmentationHo Kei Cheng, Yu-Wing Tai, Chi-Keung TangNeurIPS 2021 · 被引用 403 次
- Associating Objects with Transformers for Video Object SegmentationZongxin Yang, Yunchao Wei, Yi YangNeurIPS 2021 · 被引用 398 次
相关 Paper
- Recurrent Dynamic Embedding for Video Object SegmentationMingxing Li, Li Hu, Zhiwei Xiong, Bang Zhang 等CVPR 2022 · 被引用 80 次
- LVOS: A Benchmark for Long-term Video Object SegmentationLingyi Hong, Wenchao Chen, Zhongying Liu, Wei Zhang 等ICCV 2023 · 被引用 89 次
- Training-Free Spatio-temporal Decoupled Reasoning Video Segmentation with Adaptive Object MemoryZhengtong Zhu, Jiaqing Fan, Zhixuan Liu, Fanzhang LiAAAI 2026 · 被引用 1 次
- Dual Temporal Memory Network for Efficient Video Object SegmentationKaihua Zhang, Long Wang, Dong Liu, Bo Liu 等ACM MM 2020 · 被引用 16 次
- Look Before You Match: Instance Understanding Matters in Video Object SegmentationJunke Wang, Dongdong Chen, Zuxuan Wu, Chong Luo 等CVPR 2023
