SegEQA: Video Segmentation Based Visual Attention for Embodied Question Answering
Haonan Luo, Guosheng Lin, Zichuan Liu, Fayao Liu, Zhenmin Tang, Yazhou Yao
Abstract
Embodied Question Answering (EQA) is a newly defined research area where an agent is required to answer the user's questions by exploring the real world environment. It has attracted increasing research interests due to its broad applications in automatic driving system, in-home robots, and personal assistants. Most of the existing methods perform poorly in terms of answering and navigation accuracy due to the absence of local details and vulnerability to the ambiguity caused by complicated vision conditions. To tackle these problems, we propose a segmentation based visual attention mechanism for Embodied Question Answering. Firstly, We extract the local semantic features by introducing a novel high-speed video segmentation framework. Then by the guide of extracted semantic features, a bottom-up visual attention mechanism is proposed for the Visual Question Answering (VQA) sub-task. Further, a feature fusion strategy is proposed to guide the training of the navigator without much additional computational cost. The ablation experiments show that our method boosts the performance of VQA module by 4.2% (68.99% vs 64.73%) and leads to 3.6% (48.59% vs 44.98%) overall improvement in EQA accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2406f02-72e6-41a2-9f1d-1dbccabf3881Cited by top-tier papers4
- CRSSC: Salvage Reusable Samples from Noisy Data for Robust LearningZeren Sun, Xian-Sheng Hua, Yazhou Yao, Xiu-Shen Wei et al.ACM MM 2020 · 57 citations
- EQA-MX: Embodied Question Answering using Multimodal ExpressionMd Mofijul Islam, Alexi Gladstone, Riashat Islam, Tariq IqbalICLR 2024 · 18 citations
- Non-Salient Region Object Mining for Weakly Supervised Semantic SegmentationYazhou Yao, Tao Chen, Guo-Sen Xie, Chuanyi Zhang et al.CVPR 2021
- Jo-SRC: A Contrastive Approach for Combating Noisy LabelsYazhou Yao, Zeren Sun, Chuanyi Zhang, Fumin Shen et al.CVPR 2021
Related papers
- Env-QA: A Video Question Answering Benchmark for Comprehensive Understanding of Dynamic EnvironmentsDifei Gao, Ruiping Wang, Ziyi Bai, Xilin ChenICCV 2021 · 36 citations
- Predict Before You Explore: Predictive Planning with Specialized Memory for Embodied Question AnsweringBowen Yuan, Sisi You, Bing-Kun BaoCVPR 2026
- Language-Guided Visual Aggregation Network for Video Question AnsweringXiao Liang, Di Wang, Quan Wang, Bo Wan et al.ACM MM 2023 · 5 citations
- Multi-Factor Adaptive Vision Selection for Egocentric Video Question AnsweringHaoyu Zhang, Meng Liu, Zixin Liu, Xuemeng Song et al.ICML 2024 · 23 citations
- A Simple LLM Framework for Long-Range Video Question-AnsweringCe Zhang, Taixi Lu, Md Mohaiminul Islam, Ziyang Wang et al.EMNLP 2024 · 37 citations
