Catch Missing Details: Image Reconstruction with Frequency Augmented Variational Autoencoder
Xinmiao Lin, Yikang Li, Jenhao Hsiao, Chiuman Ho, Yu Kong
Abstract
The popular VQ-VAE models reconstruct images through learning a discrete codebook but suffer from a significant issue in the rapid quality degradation of image reconstruction as the compression rate rises. One major reason is that a higher compression rate induces more loss of visual signals on the higher frequency spectrum which reflect the details on pixel space. In this paper, a Frequency Complement Module (FCM) architecture is proposed to capture the missing frequency information for enhancing reconstruction quality. The FCM can be easily incorporated into the VQ-VAE structure, and we refer to the new model as Frequancy Augmented VAE (FA-VAE). In addition, a Dynamic Spectrum Loss (DSL) is introduced to guide the FCMs to balance between various frequencies dynamically for optimal reconstruction. FA-VAE is further extended to the text-to-image synthesis task, and a Crossattention Autoregressive Transformer (CAT) is proposed to obtain more precise semantic attributes in texts. Extensive reconstruction experiments with different compression rates are conducted on several benchmark datasets, and the results demonstrate that the proposed FA-VAE is able to restore more faithfully the details compared to SOTA methods. CAT also shows improved generation quality with better image-text semantic alignment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 662d8f95-df45-4749-8c7d-b7a56a8eb390Cited by top-tier papers10
- SGNet: Structure Guided Network via Gradient-Frequency Awareness for Depth Map Super-resolutionZhengxue Wang, Zhiqiang Yan, Jian YangAAAI 2024 · 64 citations
- FSRT: Facial Scene Representation Transformer for Face Reenactment from Factorized Appearance, Head-Pose, and Facial Expression FeaturesAndre Rochow, Max Schwarz, Sven BehnkeCVPR 2024 · 17 citations
- LG-VQ: Language-Guided Codebook LearningGuotao Liang, Baoquan Zhang, Yaowei Wang, Yunming Ye et al.NeurIPS 2024 · 14 citations
- Decompose to Understand, Fuse to Detect: Frequency-Decoupled Anomaly Detection for Encrypted Network TrafficXinglin Lian, Chengtai Cao, Ting Zhong, Yong Wang et al.INFOCOM 2026 · 8 citations
- VersVideo: Leveraging Enhanced Temporal Diffusion Models for Versatile Video GenerationJinxi Xiang, Ricong Huang, Jun Zhang, Guanbin Li et al.ICLR 2024 · 4 citations
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- Focal Frequency Loss for Image Reconstruction and SynthesisLiming Jiang, Bo Dai, Wayne Wu, Chen Change LoyICCV 2021 · 422 citations
- Towards Accurate Image Coding: Improved Autoregressive Image Generation with Dynamic Vector QuantizationMengqi Huang, Zhendong Mao, Zhuowei Chen, Yongdong ZhangCVPR 2023
- Unveiling And Addressing Dimensional Collapse In Vector Quantization Models Via Codebook RegularizationFang Zhang, Yongxin Zhu, Yihao Liu, Bin Fu et al.ICML 2026
- Rethinking the Objectives of Vector-Quantized Tokenizers for Image SynthesisYuchao Gu, Xintao Wang, Yixiao Ge, Ying Shan et al.CVPR 2024
- Towards Improved Text-Aligned Codebook Learning: Multi-Hierarchical Codebook-Text Alignment with Long TextGuotao Liang, Baoquan Zhang, Zhiyuan Wen, Junteng Zhao et al.CVPR 2025
