Learn 3D VQA Better with Active Selection and Reannotation
Shengli Zhou, Yang Liu, Feng Zheng
Abstract
3D Visual Question Answering (3D VQA) is crucial for enabling models to perceive the physical world and perform spatial reasoning. In 3D VQA, the free-form nature of answers often leads to improper annotations that can confuse or mislead models when training on the entire dataset. While other text generation tasks can mitigate this issue by learning on large-scale datasets, the scarcity of 3D scene data enlarges the negative effect of misleading annotations. Although active learning strategies can select valuable instances for training, they fail to identify and resolve misleading labels, which the oracle inevitably provides in practice. To address this issue, we propose a multi-turn interactive active learning strategy. This strategy selects data based on models' semantic uncertainty to form a solid knowledge foundation more effectively and actively requests reannotation from an oracle to resolve potentially misleading labels. For uncertainty assessment, we utilize a variance-based metric that takes semantic relationships between terms into consideration, thus avoiding the uniform inter-class similarity assumption of previous assessment metrics. Extensive experiments exhibit better model performance and a substantial reduction in training costs, with a halving of training costs for achieving relatively high accuracy. The code is available at https://github.com/fz-zsl/AQuA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2317ac26-556c-467b-8aa4-99c6b56687f2Cited by top-tier papers2
- Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language ModelsShengli Zhou, Minghang Zheng, Feng Zheng, Yang LiuCVPR 2026 · 2 citations
- Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMsWentao Mo, Yang LiuICML 2026
Builds on12
- An Embodied Generalist Agent in 3D WorldJiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu et al.ICML 2024 · 361 citations
- ScanQA: 3D Question Answering for Spatial Scene UnderstandingDaichi Azuma, Taiki Miyanishi, Shuhei Kurita, Motoaki KawanabeCVPR 2022 · 135 citations
- Influence Selection for Active LearningZhuoming Liu, Hao Ding, Huaping Zhong, Weijia Li et al.ICCV 2021 · 125 citations
- Active Contrastive Learning of Audio-Visual Video RepresentationsShuang Ma, Zhaoyang Zeng, Daniel McDuff, Yale SongICLR 2021 · 109 citations
- Uncertainty-aware Active Learning for Optimal Bayesian ClassifierGuang Zhao, Edward R. Dougherty, Byung-Jun Yoon, Francis J. Alexander et al.ICLR 2021 · 43 citations
Related papers
- AQuA: Toward Strategic Response Generation for Ambiguous Visual QuestionsJihyoung Jang, Hyounghun KimICLR 2026 · 1 citation
- AmbiRefer3D: 3D Visual Grounding with Referential AmbiguityRongjiang Zhu, Wei Kang, Zeqi Liu, Chen junyu et al.ICML 2026
- Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question AnsweringJian Lan, Zhicheng Liu, Udo Schlegel, Raoyuan Zhao et al.ICLR 2026 · 2 citations
- Semi-Supervised Active Learning with Temporal Output DiscrepancySiyu Huang, Tianyang Wang, Haoyi Xiong, Jun Huan et al.ICCV 2021 · 84 citations
- Exploring Response Uncertainty in MLLMs: An Empirical Evaluation under Misleading ScenariosYunkai Dang, Mengxi Gao, Yibo Yan, Xin Zou et al.EMNLP 2025 · 1 citation
