Gemino: Practical and Robust Neural Compression for Video Conferencing
Vibhaalakshmi Sivaraman, Pantea Karimi, Vedantha Venkatapathy, Mehrdad Khani Shirkoohi, Sadjad Fouladi, Mohammad Alizadeh, Frédo Durand, Vivienne Sze
摘要
Video conferencing systems suffer from poor user experience when network conditions deteriorate because current video codecs simply cannot operate at extremely low bitrates. Recently, several neural alternatives have been proposed that reconstruct talking head videos at very low bitrates using sparse representations of each frame such as facial landmark information. However, these approaches produce poor reconstructions in scenarios with major movement or occlusions over the course of a call, and do not scale to higher resolutions. We design Gemino, a new neural compression system for video conferencing based on a novel high-frequencyconditional super-resolution pipeline. Gemino upsamples a very low-resolution version of each target frame while enhancing high-frequency details (e.g., skin texture, hair, etc.) based on information extracted from a single high-resolution reference image. We use a multi-scale architecture that runs different components of the model at different resolutions, allowing it to scale to resolutions comparable to 720p, and we personalize the model to learn specific details of each person, achieving much better fidelity at low bitrates. We implement Gemino atop aiortc, an open-source Python implementation of WebRTC, and show that it operates on 1024×1024 videos in real-time on a Titan X GPU, and achieves 2.2-5× lower bitrate than traditional video codecs for the same perceptual quality.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- GRACE: Loss-Resilient Real-Time Video through Neural CodecsYihua Cheng, Ziyi Zhang, Hanchen Li, Anton Arapin 等NSDI 2024 · 被引用 53 次
- Fumos: Neural Compression and Progressive Refinement for Continuous Point Cloud Video StreamingZhicheng Liang, Junhua Liu, Mallesham Dasari, Fangxin WangIEEE VR 2024 · 被引用 29 次
- ACE: Sending Burstiness Control for High-Quality Real-time CommunicationXiangjie Huang, Jiayang Xu, Haiping Wang, Hebin Yu 等SIGCOMM 2025 · 被引用 8 次
- Mowgli: Passively Learned Rate Control for Real-Time VideoNeil Agarwal, Rui Pan, Francis Y. Yan, Ravi NetravaliNSDI 2025 · 被引用 7 次
- AVA: Towards Agentic Video Analytics with Vision Language ModelsYuxuan Yan, Shiqi Jiang, Ting Cao, Yifan Yang 等NSDI 2026 · 被引用 4 次
它引用的顶会 Paper6
- Neural-Enhanced Live Streaming: Improving Live Video Ingest via Online LearningJaehong Kim, Youngmok Jung, Hyunho Yeo, Juncheol Ye 等SIGCOMM 2020 · 被引用 132 次
- Efficient Video Compression via Content-Adaptive Super-ResolutionMehrdad Khani Shirkoohi, Vibhaalakshmi Sivaraman, Mohammad AlizadehICCV 2021 · 被引用 68 次
- Implicit Warping for Animation with Image SetsArun Mallya, Ting-Chun Wang, Ming-Yu LiuNeurIPS 2022 · 被引用 62 次
- One-Shot Free-View Neural Talking-Head Synthesis for Video ConferencingTing-Chun Wang, Arun Mallya, Ming-Yu LiuCVPR 2021
- Deep Face Super-Resolution With Iterative Collaboration Between Attentive Recovery and Landmark EstimationCheng Ma, Zhenyu Jiang, Yongming Rao, Jiwen Lu 等CVPR 2020
相关 Paper
- DeNC++: Efficient Diffusion-Enhanced Neural Codec for End-to-end Semantic Streaming at the EdgeQihua Zhou, Wangjiang Gong, Zili Meng, Yaxiong Xie 等AAAI 2026
- DAVD-Net: Deep Audio-Aided Video Decompression of Talking HeadsXi Zhang, Xiaolin Wu, Xinliang Zhai, Xianye Ben 等CVPR 2020
- Deep Multi-modality Soft-decoding of Very Low Bit-rate Face VideosYanhui Guo, Xi Zhang, Xiaolin WuACM MM 2020 · 被引用 9 次
- Extreme-scale Talking-Face Video Upsampling with Audio-Visual PriorsSindhu B. Hegde, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2022 · 被引用 1 次
- Random-Access Neural Compression of Material TexturesKarthik Vaidyanathan, Marco Salvi, Bartlomiej Wronski, Tomas Akenine-Möller 等SIGGRAPH 2023 · 被引用 29 次
