CompGuessWhat?!: A Multi-task Evaluation Framework for Grounded Language Learning
Alessandro Suglia, Ioannis Konstas, Andrea Vanzo, Emanuele Bastianelli, Desmond Elliott, Stella Frank, Oliver Lemon
Abstract
Approaches to Grounded Language Learning typically focus on a single task-based final performance measure that may not depend on desirable properties of the learned hidden representations, such as their ability to predict salient attributes or to generalise to unseen situations. To remedy this, we present GROLLA, an evaluation framework for Grounded Language Learning with Attributes with three subtasks: 1) Goal-oriented evaluation; 2) Object attribute prediction evaluation; and 3) Zeroshot evaluation. We also propose a new dataset CompGuessWhat?! as an instance of this framework for evaluating the quality of learned neural representations, in particular concerning attribute grounding. To this end, we extend the original GuessWhat?! dataset by including a semantic layer on top of the perceptual one. Specifically, we enrich the Vi-sualGenome scene graphs associated with the GuessWhat?! images with abstract and situated attributes. By using diagnostic classifiers, we show that current models learn representations that are not expressive enough to encode object attributes (average F1 of 44.27). In addition, they do not learn strategies nor representations that are robust enough to perform well when novel scenes or objects are involved in gameplay (zero-shot best accuracy 50.06%).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Investigating the Role of Instruction Variety and Task Difficulty in Robotic Manipulation TasksAmit Parekh, Nikolas Vitsakis, Alessandro Suglia, Ioannis KonstasEMNLP 2024 · 2 citations
- Sketchtopia: A Dataset and Foundational Agents for Benchmarking Asynchronous Multimodal Communication with Iconic FeedbackMohd Hozaifa Khan, Ravi Kiran SarvadevabhatlaCVPR 2025
Builds on3
- nocaps: novel object captioning at scaleHarsh Agrawal, Peter Anderson, Karan Desai, Yufei Wang et al.ICCV 2019 · 631 citations
- The Dialogue Dodecathlon: Open-Domain Knowledge and Image Grounded Conversational AgentsKurt Shuster, Da Ju, Stephen Roller, Emily Dinan et al.ACL 2020 · 9 citations
- 12-in-1: Multi-Task Vision and Language Representation LearningJiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh et al.CVPR 2020
Related papers
- Describe Me an Auklet: Generating Grounded Perceptual Category DescriptionsBill Noble, Nikolai IlinykhEMNLP 2023
- Learning Conditional Attributes for Compositional Zero-Shot LearningQingsheng Wang, Lingqiao Liu, Chenchen Jing, Hao Chen et al.CVPR 2023
- Disentangling Visual Embeddings for Attributes and ObjectsNirat Saini, Khoi Pham, Abhinav ShrivastavaCVPR 2022 · 74 citations
- Beyond Seen Primitive Concepts and Attribute-Object Compositional LearningNirat Saini, Khoi Pham, Abhinav ShrivastavaCVPR 2024 · 4 citations
- Zero-Shot Grounding of Objects From Natural Language QueriesArka Sadhu, Kan Chen, Ram NevatiaICCV 2019 · 176 citations
