Sim VQA: Exploring Simulated Environments for Visual Question Answering
Paola Cascante-Bonilla, Hui Wu, Letao Wang, Rogério Feris, Vicente Ordonez
摘要
Existing work on VQA explores data augmentation to achieve better generalization by perturbing images in the dataset or modifying existing questions and answers. While these methods exhibit good performance, the diversity of the questions and answers are constrained by the available images. In this work we explore using synthetic computer-generated data to fully control the visual and language space, allowing us to provide more diverse scenarios. We quantify the effectiveness of leveraging synthetic data for real-world VQA. By exploiting 3D and physics simulation platforms, we provide a pipeline to generate synthetic data to expand and replace type-specific questions and answers without risking exposure of sensitive or personal data that might be present in real images. We offer a comprehensive analysis while expanding existing hyper-realistic datasets to be usedfor VQA. We also propose Feature Swapping (F-SWAP) - where we randomly switch object-level features during training to make a VQA model more domain invariant. We show that F-SWAP is effective for improving VQA models on real images without compromising on their accuracy to answer existing questions in the dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- EarthVQA: Towards Queryable Earth via Relational Reasoning-Based Remote Sensing Visual Question AnsweringJunjue Wang, Zhuo Zheng, Zihang Chen, Ailong Ma 等AAAI 2024 · 被引用 72 次
- Going Beyond Nouns With Vision & Language Models Using Synthetic DataPaola Cascante-Bonilla, Khaled Shehada, James Seale Smith, Sivan Doveh 等ICCV 2023 · 被引用 49 次
- Beat: Bi-directional One-to-Many Embedding Alignment for Text-based Person RetrievalYiwei Ma, Xiaoshuai Sun, Jiayi Ji, Guannan Jiang 等ACM MM 2023 · 被引用 34 次
- GOI: Find 3D Gaussians of Interest with an Optimizable Open-vocabulary Semantic-space HyperplaneYansong Qu, Shaohui Dai, Xinyang Li, Jianghang Lin 等ACM MM 2024 · 被引用 20 次
- Toward Unsupervised Realistic Visual Question AnsweringYuwei Zhang, Chih-Hui Ho, Nuno VasconcelosICCV 2023 · 被引用 3 次
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong 等AAAI 2020 · 被引用 966 次
- Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene UnderstandingMike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar 等ICCV 2021 · 被引用 633 次
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli 等ICLR 2020 · 被引用 584 次
相关 Paper
- SwapMix: Diagnosing and Regularizing the Over-Reliance on Visual Context in Visual Question AnsweringVipul Gupta, Zhuowan Li, Adam Kortylewski, Chenyu Zhang 等CVPR 2022 · 被引用 41 次
- SK-VQA: Synthetic Knowledge Generation at Scale for Training Context-Augmented Multimodal LLMsXin Su, Man Luo, Kris W. Pan, Tien Pei Chou 等ICML 2025
- CrossVQA: Scalably Generating Benchmarks for Systematically Testing VQA GeneralizationArjun R. Akula, Soravit Changpinyo, Boqing Gong, Piyush Sharma 等EMNLP 2021 · 被引用 18 次
- Improving Question Answering Model Robustness with Synthetic Adversarial Data GenerationMax Bartolo, Tristan Thrush, Robin Jia, Sebastian Riedel 等EMNLP 2021 · 被引用 68 次
- Counterfactual Vision and Language LearningEhsan Abbasnejad, Damien Teney, Amin Parvaneh, Javen Shi 等CVPR 2020
