Lune

ICML2024Top-tier venue

Multi-Factor Adaptive Vision Selection for Egocentric Video Question Answering

Haoyu Zhang, Meng Liu, Zixin Liu, Xuemeng Song, Yaowei Wang, Liqiang Nie

2024Year
23Citations
15Top-tier citations

Abstract

The challenge of interpreting the world from a human perspective in Artificial Intelligence (AI) is particularly evident in egocentric video question answering, which grapples with issues like small object recognition, noise suppression, and spatial-temporal reasoning. To address these challenges, we introduce the Multi-Factor Adaptive vision Selection (MFAS) framework. MFAS integrates a patch partition and merging module for enhanced small object recognition, a priorguided patch selection module for noise suppression and focused analysis, and a hierarchical aggregation network to aggregate visual semantics guided by questions. Extensive experiments on several public egocentric datasets have validated the effectiveness and generalization of our framework. Code and data are available in https:// github.com/Hyu-Zhang/EgoVideoQA .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 51e4ea31-d2ee-464a-a14b-779f0543c9d7

Cited by top-tier papers15

Ask how each one uses it

Builds on32

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines