Pasta: Proportional Amplitude Spectrum Training Augmentation for Syn-to-Real Domain Generalization
Prithvijit Chattopadhyay, Kartik Sarangmath, Vivek Vijaykumar, Judy Hoffman
Abstract
Synthetic data offers the promise of cheap and bountiful training data for settings where labeled real-world data is scarce. However, models trained on synthetic data significantly underperform when evaluated on real-world data. In this paper, we propose Proportional Amplitude Spectrum Training Augmentation (Pasta), a simple and effective augmentation strategy to improve out-of-the-box synthetic-to-real (syn-to-real) generalization performance. Pasta perturbs the amplitude spectra of synthetic images in the Fourier domain to generate augmented views. Specifically, with Pasta we propose a structured perturbation strategy where high-frequency components are perturbed relatively more than the low-frequency ones. For the tasks of semantic segmentation (GTAV→Real), object detection (Sim10K→Real), and object recognition (VisDA-C Syn→Real), across a total of 5 syn-to-real shifts, we find that Pasta outperforms more complex state-of-the-art generalization methods while being complementary to the same.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers24
- Learning Frequency-Adapted Vision Foundation Model for Domain Generalized Semantic SegmentationQi Bi, Jingjun Yi, Hao Zheng, Haolan Zhan et al.NeurIPS 2024 · 62 citations
- Stronger, Fewer, & Superior: Harnessing Vision Foundation Models for Domain Generalized Semantic SegmentationZhixiang Wei, Lin Chen, Yi Jin, Xiaoxiao Ma et al.CVPR 2024 · 61 citations
- Exploring Semantic Consistency and Style Diversity for Domain Generalized Semantic SegmentationHongwei Niu, Linhuang Xie, Jianghang Lin, Shengchuan ZhangAAAI 2025 · 16 citations
- PiLoT: Neural Pixel-to-3D Registration for UAV-based Ego and Target Geo-localizationXiaoya Cheng, Long Wang, Yan Liu, Xinyi Liu et al.CVPR 2026 · 6 citations
- Leveraging Depth and Language for Open-Vocabulary Domain-Generalized Semantic SegmentationSiyu Chen, Ting Han, Chengzheng Fu, Changshe Zhang et al.NeurIPS 2025 · 4 citations
Builds on33
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- In Search of Lost Domain GeneralizationIshaan Gulrajani, David Lopez-PazICLR 2021 · 1,416 citations
Related papers
- Fourier-Basis Functions to Bridge Augmentation Gap: Rethinking Frequency Augmentation in Image ClassificationPuru Vaish, Shunxin Wang, Nicola StrisciuglioCVPR 2024 · 11 citations
- Do We Need All the Synthetic Data? Targeted Image Augmentation via Diffusion ModelsDang Nguyen, Jiping Li, Jinghao Zheng, Baharan MirzasoleimanICLR 2026 · 3 citations
- Generalizable Fourier Augmentation for Unsupervised Video Object SegmentationHuihui Song, Tiankang Su, Yuhui Zheng, Kaihua Zhang et al.AAAI 2024 · 15 citations
- Amplitude-Phase Recombination: Rethinking Robustness of Convolutional Neural Networks in Frequency DomainGuangyao Chen, Peixi Peng, Li Ma, Jia Li et al.ICCV 2021 · 132 citations
- VITA: A Multi-Source Vicinal Transfer Augmentation Method for Out-of-Distribution GeneralizationMinghui Chen, Cheng Wen, Feng Zheng, Fengxiang He et al.AAAI 2022 · 5 citations
