Navigating Connected Memories with a Task-oriented Dialog System
Satwik Kottur, Seungwhan Moon, Alborz Geramifard, Babak Damavandi
Abstract
Recent years have seen an increasing trend in the volume of personal media captured by users, thanks to the advent of smartphones and smart glasses, resulting in large media collections. Despite conversation being an intuitive humancomputer interface, current efforts focus mostly on single-shot natural language based media retrieval to aid users query their media and re-live their memories. This severely limits the search functionality as users can neither ask followup queries nor obtain information without first formulating a single-turn query. In this work, we propose dialogs for connected memories as a powerful tool to empower users to search their media collection through a multiturn, interactive conversation. Towards this, we collect a new task-oriented dialog dataset COMET, which contains 11.5k user↔assistant dialogs (totalling 103k utterances), grounded in simulated personal memory graphs. We employ a resource-efficient, two-phase data collection pipeline that uses: (1) a novel multimodal dialog simulator that generates synthetic dialog flows grounded in memory graphs, and, (2) manual paraphrasing to obtain natural language utterances. We analyze COMET, formulate four main tasks to benchmark meaningful progress, and adopt state-of-the-art language models as strong baselines, in order to highlight the multimodal challenges captured by our dataset 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1a41618b-4df5-4aad-85e7-a49e4875c99cCited by top-tier papers1
Ask how each one uses itBuilds on6
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Towards Scalable Multi-Domain Conversational Agents: The Schema-Guided Dialogue DatasetAbhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta et al.AAAI 2020 · 707 citations
- A Simple Language Model for Task-Oriented DialogueEhsan Hosseini-Asl, Bryan McCann, Chien-Sheng Wu, Semih Yavuz et al.NeurIPS 2020 · 590 citations
- SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal ConversationsSatwik Kottur, Seungwhan Moon, Alborz Geramifard, Babak DamavandiEMNLP 2021 · 54 citations
Related papers
- Interactive Person Retrieval via Multi-Turn Multimodal ConversationYang Bai, Tingfeng Wang, Bin Yang, Min Cao et al.ICML 2026
- OmniQuery: Contextually Augmenting Captured Multimodal Memories to Enable Personal Question AnsweringJiahao Nick Li, Zhuohao Jerry Zhang, Jiaju MaCHI 2025 · 22 citations
- Enabling Chatbots with Eyes and Ears: An Immersive Multimodal Conversation System for Dynamic InteractionsJihyoung Jang, Minwook Bae, Minji Kim, Dilek Hakkani-Tür et al.ACL 2025
- TikTalk: A Video-Based Dialogue Dataset for Multi-Modal Chitchat in Real WorldHongpeng Lin, Ludan Ruan, Wenke Xia, Peiyu Liu et al.ACM MM 2023 · 10 citations
- MuCo: Multi-turn Contrastive Learning for Multimodal Embedding ModelGeonmo Gu, Byeongho Heo, Jaemyung Yu, Jaehui Hwang et al.CVPR 2026 · 2 citations
