Lune

CVPR2025Top-tier venue

Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos

Chiara Plizzari, Alessio Tonioni, Yongqin Xian, Achin Kulshrestha, Federico Tombari

2025Year
13Top-tier citations

Abstract

Our dataset: EgoTempo Free-form Video Q&A Temporal Event Ordering Q: What does the person do after draining the excess water? Object Counting Q: What is the sequence of actions the person performs with the mug? Multi-Modal LLMs Temporal Understanding Limitations of previous egocentric VideoQA datasets Single-frame Understanding Commonsense Reasoning Q: What is the main purpose of using aluminum foil? Q: What is the status of the microwave before the user gets something from it? Action Sequence Q: How many oranges does the person pick from the tree?

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext c9842f08-db09-48ae-9815-66dd58f683fb

Cited by top-tier papers13

Ask how each one uses it

Builds on18

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines