Gaussian Mixture Proposals with Pull-Push Learning Scheme to Capture Diverse Events for Weakly Supervised Temporal Video Grounding
Sunoh Kim, Jungchan Cho, Joonsang Yu, Youngjoon Yoo, Jin Young Choi
摘要
In the weakly supervised temporal video grounding study, previous methods use predetermined single Gaussian proposals which lack the ability to express diverse events described by the sentence query. To enhance the expression ability of a proposal, we propose a Gaussian mixture proposal (GMP) that can depict arbitrary shapes by learning importance, centroid, and range of every Gaussian in the mixture. In learning GMP, each Gaussian is not trained in a feature space but is implemented over a temporal location. Thus the conventional feature-based learning for Gaussian mixture model is not valid for our case. In our special setting, to learn moderately coupled Gaussian mixture capturing diverse events, we newly propose a pull-push learning scheme using pulling and pushing losses, each of which plays an opposite role to the other. The effects of components in our scheme are verified in-depth with extensive ablation studies and the overall scheme achieves state-of-the-art performance. Our code is available at https://github.com/sunoh-kim/pps .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video CaptioningShiping Ge, Qiang Chen, Zhiwei Jiang, Yafeng Yin 等AAAI 2025 · 被引用 7 次
- When CLIP Sees More, It Fights Back Harder: Multi-View Guided Adaptive Counterattacks for Test-Time Adversarial RobustnessSunoh Kim, Daeho UmCVPR 2026 · 被引用 2 次
- PC-Net: Weakly Supervised Compositional Moment Retrieval via Proposal-Centric NetworkMingyao Zhou, Hao Sun, Wei Xie, Ming Dong 等NeurIPS 2025 · 被引用 1 次
- TubeRMC: Tube-conditioned Reconstruction with Mutual Constraints for Weakly-supervised Spatio-Temporal Video GroundingJinxuan Li, Yi Zhang, Jian-Fang Hu, Chaolei Tan 等AAAI 2026 · 被引用 1 次
- Boundary-Aware Temporal Dynamic Pseudo-Supervision Pairs Generation for Zero-Shot Natural Language Video LocalizationXiongwen Deng, Haoyu Tang, Han Jiang, Qinghai Zheng 等AAAI 2025
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Weakly-Supervised Video Moment Retrieval via Semantic Completion NetworkZhijie Lin, Zhou Zhao, Zhu Zhang, Qi Wang 等AAAI 2020 · 被引用 170 次
- Weakly Supervised Video Moment Localization with Contrastive Negative Sample MiningMinghang Zheng, Yanjie Huang, Qingchao Chen, Yang LiuAAAI 2022 · 被引用 109 次
- Weakly Supervised Temporal Sentence Grounding with Gaussian-based Contrastive Proposal LearningMinghang Zheng, Yanjie Huang, Qingchao Chen, Yuxin Peng 等CVPR 2022 · 被引用 108 次
- Regularized Two-Branch Proposal Networks for Weakly-Supervised Moment Retrieval in VideosZhu Zhang, Zhijie Lin, Zhou Zhao, Jieming Zhu 等ACM MM 2020 · 被引用 86 次
相关 Paper
- D3G: Exploring Gaussian Prior for Temporal Sentence Grounding with Glance AnnotationHanjun Li, Xiujun Shu, Sunan He, Ruizhi Qiao 等ICCV 2023 · 被引用 21 次
- Activity-driven Weakly-Supervised Spatio-Temporal Grounding from Untrimmed VideosJunwen Chen, Wentao Bao, Yu KongACM MM 2020 · 被引用 19 次
- Video-Text Prompting for Weakly Supervised Spatio-Temporal Video GroundingHeng Zhao, Yinjie Zhao, Bihan Wen, Yew-Soon Ong 等EMNLP 2024
- Explore Inter-contrast between Videos via Composition for Weakly Supervised Temporal Sentence GroundingJiaming Chen, Weixin Luo, Wei Zhang, Lin MaAAAI 2022 · 被引用 33 次
- Unsupervised Temporal Video Grounding with Deep Semantic ClusteringDaizong Liu, Xiaoye Qu, Yinzhen Wang, Xing Di 等AAAI 2022 · 被引用 52 次
