Zero-Shot Compositional Video Learning with Coding Rate Reduction
Heeseok Jung, Jun-Hyeon Bak, Yujin Jeong, Gyugeun Lee, Jinwoo Ahn, Eun-Sol Kim
摘要
In this paper, we propose a novel zero-shot compositional video understanding method inspired by how young children efficiently learn new concepts and flexibly expand their existing knowledge framework. While recent large-scale visual language models (VLMs) have achieved remarkable advancements and demonstrated impressive performance improvements across various tasks, they require massive amounts of data and computational resources. However, despite their strong benchmark performance, they often fail to solve simple zero-shot composition tasks. Moreover, VLMs designed for video data demand even greater computational resources. We introduce a new video representation learning method inspired by human compositional learning to address these challenges. Specifically, we demonstrate that achieving zero-shot compositional learning requires effective representation learning that disentangles given data into meaningful semantic units. We propose a novel method that learns such disentangled representations based on an information-theoretic measure. By optimizing coding rate reduction, we successfully learn spatio-temporally disentangled features from videos, one of the most challenging data. Our approach significantly enhances compositional generalizability, demonstrating its effectiveness in zero-shot learning scenarios. Codes will be available at https://github.com/heeseokjung/zs-comp-mcr2.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper35
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- Learning Visual Proxy for Compositional Zero-Shot LearningShiyu Zhang, Cheng Yan, Yang Liu, Chenchen Jing 等ICCV 2025 · 被引用 1 次
- FlowComposer: Composable Flows for Compositional Zero-Shot LearningZhenqi He, Lin Li, Long ChenCVPR 2026 · 被引用 3 次
- Decomposed Soft Prompt Guided Fusion Enhancing for Compositional Zero-Shot LearningXiaocheng Lu, Song Guo, Ziming Liu, Jingcai GuoCVPR 2023
- Compositional Video Understanding with Spatiotemporal Structure-based TransformersHoyeoung Yun, Jinwoo Ahn, Minseo Kim, Eun-Sol KimCVPR 2024 · 被引用 4 次
- Orthogonal Temporal Interpolation for Zero-Shot Video RecognitionYan Zhu, Junbao Zhuo, Bin Ma, Jiajia Geng 等ACM MM 2023 · 被引用 5 次
