Morphe: High-Fidelity Generative Video Streaming with Vision Foundation Model
Tianyi Gong, Zijian Cao, Zixing Zhang, Jiangkai Wu, Xinggong Zhang, Shuguang Cui, Fangxin Wang
Abstract
Video streaming is a fundamental Internet service, while the quality still cannot be guaranteed especially in poor network conditions such as bandwidth-constrained and remote areas. Existing works mainly work towards two directions: traditional pixel-codec streaming nearly approaches its limit and is hard to step further in compression; the emerging neuralenhanced or generative streaming usually fall short in latency and visual fidelity, hindering their practical deployment.
Inspired by the recent success of vision foundation model (VFM), we strive to harness the powerful video understanding and processing capacities of VFM to achieve generalization, high fidelity and loss resilience for real-time video streaming with even higher compression rate. We present Morphe 1 , the first revolutionized paradigm that enables VFM-based endto-end generative video streaming towards this goal. Specifically, Morphe employs joint training of visual tokenizers and variable-resolution spatiotemporal optimization under simulated network constraints. Additionally, a robust streaming system is constructed that leverages intelligent packet dropping to resist real-world network perturbations. Extensive evaluation demonstrates that Morphe achieves comparable visual quality while saving 62.5% bandwidth compared to H.265, and accomplishes real-time, loss-resilient video delivery in challenging network environments, representing a milestone in VFM-enabled multimedia streaming solutions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on15
- Learning in situ: a randomized experiment in video streamingFrancis Y. Yan, Hudson Ayers, Chenzhi Zhu, Sadjad Fouladi et al.NSDI 2020 · 360 citations
- Neural Inter-Frame Compression for Video CodingAbdelaziz Djelouah, Joaquim Campos, Simone Schaub-Meyer, Christopher SchroersICCV 2019 · 207 citations
- Neural-Enhanced Live Streaming: Improving Live Video Ingest via Online LearningJaehong Kim, Youngmok Jung, Hyunho Yeo, Juncheol Ye et al.SIGCOMM 2020 · 132 citations
- NEMO: enabling neural-enhanced video streaming on commodity mobile devicesHyunho Yeo, Chan Ju Chong, Youngmok Jung, Juncheol Ye et al.MobiCom 2020 · 118 citations
- Real-world Video Super-resolution: A Benchmark Dataset and A Decomposition based Learning SchemeXi Yang, Wangmeng Xiang, Hui Zeng, Lei ZhangICCV 2021 · 90 citations
Related papers
- Evaluating Newtonian Mechanics in Video Generative Models with Real Physical SystemsAntonios Tragoudaras, Chenyu Zhang, Daniil Cherniavskii, Antonis Vozikis et al.ICML 2026 · 39 citations
- VFMF: Dense Forecasting by Generating Foundation Model FeaturesGabrijel Boduljak, Yushi Lan, Christian Rupprecht, Andrea VedaldiICML 2026
- High-Quality Joint Image and Video Tokenization with Causal VAEDawit Mureja Argaw, Xian Liu, Qinsheng Zhang, Joon Son Chung et al.ICLR 2025
- DeNC: Unleash Neural Codecs in Video Streaming with Diffusion EnhancementQihua Zhou, Ruibin Li, Jingcai Guo, Yaodong Huang et al.AAAI 2025 · 2 citations
- Learned Video CompressionOren Rippel, Sanjay Nair, Carissa Lew, Steve Branson et al.ICCV 2019 · 258 citations
