Hiding Video in Audio via Reversible Generative Models
Hyukryul Yang, Hao Ouyang, Vladlen Koltun, Qifeng Chen
Abstract
We present a method for hiding video content inside audio files while preserving the perceptual fidelity of the cover audio. This is a form of cross-modal steganography and is particularly challenging due to the high bitrate of video. Our scheme uses recent advances in flow-based generative models, which enable mapping audio to latent codes such that nearby codes correspond to perceptually similar signals. We show that compressed video data can be concealed in the latent codes of audio sequences while preserving the fidelity of both the hidden video and the cover audio. We can embed 128x128 video inside same-duration audio, or higher-resolution video inside longer audio sequences. Quantitative experiments show that our approach outperforms relevant baselines in steganographic capacity and fidelity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3fb695c5-5f7b-438d-a49f-50cff6a8e3e5Cited by top-tier papers5
- StegaNeRF: Embedding Invisible Information within Neural Radiance FieldsChenxin Li, Brandon Y. Feng, Zhiwen Fan, Panwang Pan et al.ICCV 2023 · 57 citations
- IICNet: A Generic Framework for Reversible Image ConversionKa Leong Cheng, Yueqi Xie, Qifeng ChenICCV 2021 · 30 citations
- Images that Sound: Composing Images and Sounds on a Single CanvasZiyang Chen, Daniel Geng, Andrew OwensNeurIPS 2024 · 22 citations
- Embedding Novel Views in a Single JPEG ImageYue Wu, Guotao Meng, Qifeng ChenICCV 2021 · 16 citations
- Restorable Image Operators with Quasi-Invertible NetworksHao Ouyang, Tengfei Wang, Qifeng ChenAAAI 2022 · 7 citations
Related papers
- MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video GenerationMingzhen Sun, Weining Wang, Yanyuan Qiao, Jiahui Sun et al.ACM MM 2024 · 4 citations
- VAFlow: Video-to-Audio Generation with Cross-Modality Flow MatchingXihua Wang, Xin Cheng, Yuyue Wang, Ruihua Song et al.ICCV 2025 · 6 citations
- Aligning What Matters: Masked Latent Adaptation for Text-to-Audio-Video GenerationJiyang Zheng, Siqi Pan, Yu Yao, Zhaoqing Wang et al.NeurIPS 2025 · 6 citations
- Leveraging Real Talking Faces via Self-Supervision for Robust Forgery DetectionAlexandros Haliassos, Rodrigo Mira, Stavros Petridis, Maja PanticCVPR 2022 · 138 citations
- AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video GenerationMoayed Haji-Ali, Willi Menapace, Aliaksandr Siarohin, Ivan Skorokhodov et al.ICCV 2025 · 3 citations
