PatchGame: Learning to Signal Mid-level Patches in Referential Games
Kamal Gupta, Gowthami Somepalli, Anubhav Gupta, Vinoj Yasanga Jayasundara Magalle Hewa, Matthias Zwicker, Abhinav Shrivastava
Abstract
We study a referential game (a type of signaling game) where two agents communicate with each other via a discrete bottleneck to achieve a common goal. In our referential game, the goal of the speaker is to compose a message or a symbolic representation of "important" image patches, while the task for the listener is to match the speaker's message to a different view of the same image. We show that it is indeed possible for the two agents to develop a communication protocol without explicit or implicit supervision. We further investigate the developed protocol and show the applications in speeding up recent Vision Transformers by using only important patches, and as pre-training for downstream recognition tasks (e.g., classification). Code is available at https://kampta.github.io/patch-game .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Brain Decodes Deep NetsHuzheng Yang, James C. Gee, Jianbo ShiCVPR 2024 · 10 citations
- Semantics and Spatiality of Emergent CommunicationRotem Ben Zion, Boaz Carmeli, Orr Paradise, Yonatan BelinkovNeurIPS 2024 · 5 citations
Builds on16
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- Learning Multi-Object Positional Relationships via Emergent CommunicationYicheng Feng, Boshi An, Zongqing LuAAAI 2024 · 4 citations
- Interpretable agent communication from scratch (with a generic visual processor emerging on the side)Roberto Dessì, Eugene Kharitonov, Marco BaroniNeurIPS 2021 · 33 citations
- Emergent Communication of GeneralizationsJesse Mu, Noah D. GoodmanNeurIPS 2021 · 60 citations
- On the interaction between supervision and self-play in emergent communicationRyan Lowe, Abhinav Gupta, Jakob N. Foerster, Douwe Kiela et al.ICLR 2020 · 30 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
