Non-Semantics Suppressed Mask Learning for Unsupervised Video Semantic Compression
Yuan Tian, Guo Lu, Guangtao Zhai, Zhiyong Gao
Abstract
Most video compression methods aim to improve the decoded video visual quality, instead of particularly guaranteeing the semantic-completeness, which deteriorates downstream video analysis tasks, e.g., action recognition. In this paper, we focus on a novel unsupervised video semantic compression problem, where video semantics is compressed in a downstream task-agnostic manner. To tackle this problem, we first propose a Semantic-Mining-then-Compensation (SMC) framework to enhance the plain video codec with powerful semantic coding capability. Then, we optimize the framework with only unlabeled video data, by masking out a proportion of the compressed video and reconstructing the masked regions of the original video, which is inspired by recent masked image modeling (MIM) methods. Although the MIM scheme learns generalizable semantic features, its inner generative learning paradigm may also facilitate the coding framework memorizing non-semantic information with extra bit costs. To suppress this deficiency, we explicitly decrease the non-semantic information entropy of the decoded video features, by formulating it as a parametrized Gaussian Mixture Model conditioned on the mined video semantics. Comprehensive experimental results demonstrate the proposed approach shows remarkable superiority over previous traditional, learnable, and perceptual quality-oriented video codecs, on three video analysis tasks and seven datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9282e343-e2c7-4a51-b44b-d47d80083a4dCited by top-tier papers7
- Medical Manifestation-Aware De-IdentificationYuan Tian, Shuo Wang, Guangtao ZhaiAAAI 2025 · 7 citations
- Feature Coding in the Era of Large Models: Dataset, Test Conditions, and BenchmarkChangsheng Gao, Yifan Ma, Qiaoxi Chen, Yenan Xu et al.ICCV 2025 · 3 citations
- Semantics Versus Identity: A Divide-and-Conquer Approach Towards Adjustable Medical Image De-IdentificationYuan Tian, Shuo Wang, Rongzhao Zhang, Zijian Chen et al.ICCV 2025 · 3 citations
- DT-UFC: Universal Large Model Feature Coding via Peaky-to-Balanced Distribution TransformationChangsheng Gao, Zijie Liu, Li Li, Dong Liu et al.ACM MM 2025 · 2 citations
- Task-Aware Encoder Control for Deep Video CompressionXingtong Ge, Jixiang Luo, Xinjie Zhang, Tongda Xu et al.CVPR 2024
Builds on31
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- What Makes for Good Views for Contrastive Learning?Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan et al.NeurIPS 2020 · 1,631 citations
Related papers
- Video Compression With Rate-Distortion AutoencodersAmirHossein Habibian, Ties van Rozendaal, Jakub M. Tomczak, Taco CohenICCV 2019 · 233 citations
- Self-Conditioned Probabilistic Learning of Video RescalingYuan Tian, Guo Lu, Xiongkuo Min, Zhaohui Che et al.ICCV 2021 · 38 citations
- Cross Modal Compression: Towards Human-comprehensible Semantic CompressionJiguo Li, Chuanmin Jia, Xinfeng Zhang, Siwei Ma et al.ACM MM 2021 · 24 citations
- Unsupervised Action Segmentation via Fast Learning of Semantically Consistent ActomsZheng Xing, Weibing ZhaoAAAI 2024 · 18 citations
- Exploiting Motion Information from Unlabeled Videos for Static Image Action RecognitionYiyi Zhang, Li Niu, Ziqi Pan, Meichao Luo et al.AAAI 2020 · 7 citations
