Gemino: Practical and Robust Neural Compression for Video Conferencing
Vibhaalakshmi Sivaraman, Pantea Karimi, Vedantha Venkatapathy, Mehrdad Khani Shirkoohi, Sadjad Fouladi, Mohammad Alizadeh, Frédo Durand, Vivienne Sze
Abstract
Video conferencing systems suffer from poor user experience when network conditions deteriorate because current video codecs simply cannot operate at extremely low bitrates. Recently, several neural alternatives have been proposed that reconstruct talking head videos at very low bitrates using sparse representations of each frame such as facial landmark information. However, these approaches produce poor reconstructions in scenarios with major movement or occlusions over the course of a call, and do not scale to higher resolutions. We design Gemino, a new neural compression system for video conferencing based on a novel high-frequencyconditional super-resolution pipeline. Gemino upsamples a very low-resolution version of each target frame while enhancing high-frequency details (e.g., skin texture, hair, etc.) based on information extracted from a single high-resolution reference image. We use a multi-scale architecture that runs different components of the model at different resolutions, allowing it to scale to resolutions comparable to 720p, and we personalize the model to learn specific details of each person, achieving much better fidelity at low bitrates. We implement Gemino atop aiortc, an open-source Python implementation of WebRTC, and show that it operates on 1024×1024 videos in real-time on a Titan X GPU, and achieves 2.2-5× lower bitrate than traditional video codecs for the same perceptual quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- GRACE: Loss-Resilient Real-Time Video through Neural CodecsYihua Cheng, Ziyi Zhang, Hanchen Li, Anton Arapin et al.NSDI 2024 · 53 citations
- Fumos: Neural Compression and Progressive Refinement for Continuous Point Cloud Video StreamingZhicheng Liang, Junhua Liu, Mallesham Dasari, Fangxin WangIEEE VR 2024 · 29 citations
- ACE: Sending Burstiness Control for High-Quality Real-time CommunicationXiangjie Huang, Jiayang Xu, Haiping Wang, Hebin Yu et al.SIGCOMM 2025 · 8 citations
- Mowgli: Passively Learned Rate Control for Real-Time VideoNeil Agarwal, Rui Pan, Francis Y. Yan, Ravi NetravaliNSDI 2025 · 7 citations
- AVA: Towards Agentic Video Analytics with Vision Language ModelsYuxuan Yan, Shiqi Jiang, Ting Cao, Yifan Yang et al.NSDI 2026 · 4 citations
Builds on6
- Neural-Enhanced Live Streaming: Improving Live Video Ingest via Online LearningJaehong Kim, Youngmok Jung, Hyunho Yeo, Juncheol Ye et al.SIGCOMM 2020 · 132 citations
- Efficient Video Compression via Content-Adaptive Super-ResolutionMehrdad Khani Shirkoohi, Vibhaalakshmi Sivaraman, Mohammad AlizadehICCV 2021 · 68 citations
- Implicit Warping for Animation with Image SetsArun Mallya, Ting-Chun Wang, Ming-Yu LiuNeurIPS 2022 · 62 citations
- One-Shot Free-View Neural Talking-Head Synthesis for Video ConferencingTing-Chun Wang, Arun Mallya, Ming-Yu LiuCVPR 2021
- Deep Face Super-Resolution With Iterative Collaboration Between Attentive Recovery and Landmark EstimationCheng Ma, Zhenyu Jiang, Yongming Rao, Jiwen Lu et al.CVPR 2020
Related papers
- DeNC++: Efficient Diffusion-Enhanced Neural Codec for End-to-end Semantic Streaming at the EdgeQihua Zhou, Wangjiang Gong, Zili Meng, Yaxiong Xie et al.AAAI 2026
- DAVD-Net: Deep Audio-Aided Video Decompression of Talking HeadsXi Zhang, Xiaolin Wu, Xinliang Zhai, Xianye Ben et al.CVPR 2020
- Deep Multi-modality Soft-decoding of Very Low Bit-rate Face VideosYanhui Guo, Xi Zhang, Xiaolin WuACM MM 2020 · 9 citations
- Extreme-scale Talking-Face Video Upsampling with Audio-Visual PriorsSindhu B. Hegde, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2022 · 1 citation
- Random-Access Neural Compression of Material TexturesKarthik Vaidyanathan, Marco Salvi, Bartlomiej Wronski, Tomas Akenine-Möller et al.SIGGRAPH 2023 · 29 citations
