Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
Apratim Bhattacharyya, Bicheng Xu, Sanjay Haresh, Reza Pourreza, Litian Liu, Sunny Panchal, Leonid Sigal, Roland Memisevic
Abstract
Multi-modal Large Language Models (LLM) have advanced conversational abilities but struggle with providing live, interactive step-by-step guidance, a key capability for future AI assistants. Effective guidance requires not only delivering instructions but also detecting their successful execution, as well as identifying and alerting users to mistakes, all of which has to happen in real-time. This requires models that are not turn-based, but that can react asynchronously to a video stream, as well as video data showing users performing tasks including mistakes and their corrections. To this end, we introduce Qualcomm Interactive Cooking, a new benchmark and dataset built upon CaptainCook4D, which contains user mistakes during task execution. Our dataset and benchmark features densely annotated, timed instructions and feedback messages, specifically including mistake alerts precisely timestamped to their visual occurrence in the video. We evaluate state-ofthe-art multi-modal LLMs on the Qualcomm Interactive Cooking benchmark and introduce LIVEMAMBA, a streaming multi-modal LLM designed for interactive instructional guidance. This work provides the first dedicated benchmark and a strong baseline for developing and evaluating on live, situated coaching.
To address the challenge of live, step-by-step coaching, we introduce the Qualcomm Interactive Cooking dataset and benchmark, as currently available large-scale vision-language datasets and * Equal contribution. † Work done while employed at Qualcomm AI Research. ‡ Qualcomm AI Research is an initiative of Qualcomm Technologies, Inc. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
1: Slice one tomato into about 1/2 inch thick slices. 2: Place the thick slices of tomatoes on a platter, ensuring they only make a single layer. 3: Season the tomato slices with salt. 4: Season the platter with 1/4 teaspoon of black pepper. 5: Sprinkle mozzarella cheese on top of the tomato throughout the platter. 6: Garnish the platter with Italian seasoning. 7: Add a drizzle of extra-virgin olive oil, about 1 tablespoon, over the entire platter.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a58baf40-701c-4158-8887-b8ed68df334cCited by top-tier papers1
Ask how each one uses itBuilds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- LiViBench: An Omnimodal Benchmark for Interactive Livestream Video UnderstandingXiaodong Wang, Langling Huang, Zhirong Wu, Xu Zhao et al.AAAI 2026 · 1 citation
- A Highly Clean Recipe Dataset with Ingredient States Annotation for State Probing TaskMashiro Toyooka, Kiyoharu Aizawa, Yoko YamakataACM MM 2025
- "Mango Mango, How to Let The Lettuce Dry Without A Spinner?": Exploring User Perceptions of Using An LLM-Based Conversational Assistant Toward Cooking PartnerSzeyi Chan, Jiachen Li, Bingsheng Yao, Amama Mahmood et al.CSCW 2025 · 4 citations
- Can Vision-Language Models Answer Face to Face Questions in the Real-World?Reza Pourreza, Rishit Dagli, Apratim Bhattacharyya, Sunny Panchal et al.ICLR 2026 · 7 citations
- RIVER: A Real-Time Interaction Benchmark for Video LLMsYansong Shi, Qingsong Zhao, Tianxiang Jiang, Xiangyu Zeng et al.ICLR 2026 · 12 citations
