Non-Semantics Suppressed Mask Learning for Unsupervised Video Semantic Compression
Yuan Tian, Guo Lu, Guangtao Zhai, Zhiyong Gao
摘要
Most video compression methods aim to improve the decoded video visual quality, instead of particularly guaranteeing the semantic-completeness, which deteriorates downstream video analysis tasks, e.g., action recognition. In this paper, we focus on a novel unsupervised video semantic compression problem, where video semantics is compressed in a downstream task-agnostic manner. To tackle this problem, we first propose a Semantic-Mining-then-Compensation (SMC) framework to enhance the plain video codec with powerful semantic coding capability. Then, we optimize the framework with only unlabeled video data, by masking out a proportion of the compressed video and reconstructing the masked regions of the original video, which is inspired by recent masked image modeling (MIM) methods. Although the MIM scheme learns generalizable semantic features, its inner generative learning paradigm may also facilitate the coding framework memorizing non-semantic information with extra bit costs. To suppress this deficiency, we explicitly decrease the non-semantic information entropy of the decoded video features, by formulating it as a parametrized Gaussian Mixture Model conditioned on the mined video semantics. Comprehensive experimental results demonstrate the proposed approach shows remarkable superiority over previous traditional, learnable, and perceptual quality-oriented video codecs, on three video analysis tasks and seven datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Medical Manifestation-Aware De-IdentificationYuan Tian, Shuo Wang, Guangtao ZhaiAAAI 2025 · 被引用 7 次
- Feature Coding in the Era of Large Models: Dataset, Test Conditions, and BenchmarkChangsheng Gao, Yifan Ma, Qiaoxi Chen, Yenan Xu 等ICCV 2025 · 被引用 3 次
- Semantics Versus Identity: A Divide-and-Conquer Approach Towards Adjustable Medical Image De-IdentificationYuan Tian, Shuo Wang, Rongzhao Zhang, Zijian Chen 等ICCV 2025 · 被引用 3 次
- DT-UFC: Universal Large Model Feature Coding via Peaky-to-Balanced Distribution TransformationChangsheng Gao, Zijie Liu, Li Li, Dong Liu 等ACM MM 2025 · 被引用 2 次
- Task-Aware Encoder Control for Deep Video CompressionXingtong Ge, Jixiang Luo, Xinjie Zhang, Tongda Xu 等CVPR 2024
它引用的顶会 Paper31
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- What Makes for Good Views for Contrastive Learning?Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan 等NeurIPS 2020 · 被引用 1,631 次
相关 Paper
- Video Compression With Rate-Distortion AutoencodersAmirHossein Habibian, Ties van Rozendaal, Jakub M. Tomczak, Taco CohenICCV 2019 · 被引用 233 次
- Self-Conditioned Probabilistic Learning of Video RescalingYuan Tian, Guo Lu, Xiongkuo Min, Zhaohui Che 等ICCV 2021 · 被引用 38 次
- Cross Modal Compression: Towards Human-comprehensible Semantic CompressionJiguo Li, Chuanmin Jia, Xinfeng Zhang, Siwei Ma 等ACM MM 2021 · 被引用 24 次
- Unsupervised Action Segmentation via Fast Learning of Semantically Consistent ActomsZheng Xing, Weibing ZhaoAAAI 2024 · 被引用 18 次
- Exploiting Motion Information from Unlabeled Videos for Static Image Action RecognitionYiyi Zhang, Li Niu, Ziqi Pan, Meichao Luo 等AAAI 2020 · 被引用 7 次
