PQA: Perceptual Question Answering
Yonggang Qi, Kai Zhang, Aneeshan Sain, Yi-Zhe Song
Abstract
Perceptual organization remains one of the very few established theories on the human visual system. It underpinned many pre-deep seminal works on segmentation and detection, yet research has seen a rapid decline since the preferential shift to learning deep models. Of the limited attempts, most aimed at interpreting complex visual scenes using perceptual organizational rules. This has however been proven to be sub-optimal, since models were unable to effectively capture the visual complexity in real-world imagery. In this paper, we rejuvenate the study of perceptual organization, by advocating two positional changes: (i) we examine purposefully generated synthetic data, instead of complex real imagery, and (ii) we ask machines to synthesize novel perceptually-valid patterns, instead of explaining existing data. Our overall answer lies with the introduction of a novel visual challenge -the challenge of perceptual question answering (PQA). Upon observing example perceptual question-answer pairs, the goal for PQA is to solve similar questions by generating answers entirely from scratch (see Figure 1 ). Our first contribution is therefore the first dataset of perceptual question-answer pairs, each generated specifically for a particular Gestalt principle. We then borrow insights from human psychology to design an agent that casts perceptual organization as a self-attention problem, where a proposed grid-to-grid mapping network directly generates answer patterns from scratch. Experiments show our agent to outperform a selection of naive and strong baselines. A human study however indicates that ours uses astronomically more data to learn when compared to an average human, necessitating future research (with or without our dataset).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Bongard-RWR+: Real-World Representations of Fine-Grained Concepts in Bongard ProblemsSzymon Pawlonka, Mikołaj Małkiński, Jacek MańdziukICLR 2026 · 7 citations
- ReasonVQA: A Multi-Hop Reasoning Benchmark with Structural Knowledge for Visual Question AnsweringDuong T. Tran, Trung-Kien Tran, Manfred Hauswirth, Danh Le PhuocICCV 2025 · 2 citations
- Reasoning Limitations of Multimodal Large Language Models. A case study of Bongard ProblemsMikolaj Malkinski, Szymon Pawlonka, Jacek MandziukICML 2025
- Unicode Analogies: An Anti-Objectivist Visual Reasoning ChallengeSteven Spratley, Krista A. Ehinger, Tim MillerCVPR 2023
Builds on2
Related papers
- Perceptual Score: What Data Modalities Does Your Model Perceive?Itai Gat, Idan Schwartz, Alexander G. SchwingNeurIPS 2021 · 56 citations
- Exploring Figure-Ground Assignment Mechanism in Perceptual OrganizationWei Zhai, Yang Cao, Jing Zhang, Zheng-Jun ZhaNeurIPS 2022 · 36 citations
- Learning to Answer Questions in Dynamic Audio-Visual ScenariosGuangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu et al.CVPR 2022 · 101 citations
- Progressive Spatio-temporal Perception for Audio-Visual Question AnsweringGuangyao Li, Wenxuan Hou, Di HuACM MM 2023 · 39 citations
- Perception Matters: Detecting Perception Failures of VQA Models Using Metamorphic TestingYuanyuan Yuan, Shuai Wang, Mingyue Jiang, Tsong Yueh ChenCVPR 2021
