MOSCATO: Predicting Multiple Object State Change through Actions
Parnian Zameni, Yuhan Shen, Ehsan Elhamifar
摘要
We introduce MOSCATO: a new benchmark for predicting the evolving states of multiple objects through long procedural videos with multiple actions. While prior work in object state prediction has typically focused on a single object undergoing one or a few state changes, realworld tasks require tracking many objects whose states evolve over multiple actions. Given the high cost of gathering framewise object-state labels for many videos, we develop a weakly-supervised multiple object state prediction framework, which only uses action labels during training. Specifically, we propose a novel Pseudo-Label Acquisition (PLA) pipeline that integrates large language models, vision-language models, and action segment annotations to generate fine-grained, per-frame object-state pseudo-labels for training a Multiple Object State Prediction (MOSP) network. We further devise a State-Action Interaction (SAI) module that explicitly models the correlations between actions and object states, thereby improving MOSP. To facilitate comprehensive evaluation, we create the MOSCATO benchmark by augmenting three egocentric video datasets with framewise object-state annotations. Experiments show that our multi-stage pseudo-labeling approach and SAI module significantly boost performance over zero-shot VLM baselines and naive extensions of existing methods, underscoring the importance of holistic action-state modeling for fine-grained procedural video understanding. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Error Recognition in Procedural Videos Using Generalized Task GraphShih-Po Lee, Ehsan ElhamifarICCV 2025 · 被引用 3 次
- AXG-Reasoner: Error Detection and Explanation in Long Task Videos with Vision–Language ModelsShih-Po Lee, Ehsan ElhamifarCVPR 2026 · 被引用 3 次
- HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction DynamicsMasatoshi Tateno, Gido Kato, Hirokatsu Kataoka, Yoichi Sato 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
- Task-Driven Modular Networks for Zero-Shot Compositional LearningSenthil Purushwalkam, Maximilian Nickel, Abhinav Gupta, Marc'Aurelio RanzatoICCV 2019 · 被引用 222 次
- VisualGPT: Data-efficient Adaptation of Pretrained Language Models for Image CaptioningJun Chen, Han Guo, Kai Yi, Boyang Li 等CVPR 2022 · 被引用 169 次
相关 Paper
- Video State-Changing Object SegmentationJiangwei Yu, Xiang Li, Xinran Zhao, Hongming Zhang 等ICCV 2023 · 被引用 16 次
- OSCBench: Benchmarking Object State Change in Text-to-Video GenerationXianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li 等ACL 2026 · 被引用 2 次
- Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture AssemblyAditya Chetan, Eric Cai, Peeyush Kushwaha, Bharath Raj Nagoor Kani 等CVPR 2026
- SAGE: A Unified Framework for Generalizable Object State Recognition with State-Action Graph EmbeddingYuan Zang, Zitian Tang, Junho Cho, Jaewook Yoo 等NeurIPS 2025
- Learning Object State Changes in Videos: An Open-World PerspectiveZihui Xue, Kumar Ashutosh, Kristen GraumanCVPR 2024 · 被引用 12 次
