Lune

ICCV2019Top-tier venue

Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning

Jyoti Aneja, Harsh Agrawal, Dhruv Batra, Alexander G. Schwing

2019Year
71Citations
22Top-tier citations

Abstract

Diverse and accurate vision+language modeling is an important goal to retain creative freedom and maintain user engagement. However, adequately capturing the intricacies of diversity in language models is challenging. Recent works commonly resort to latent variable models augmented with more or less supervision from object detectors or part-of-speech tags [10, 40] . Common to all those methods is the fact that the latent variable either only initializes the sentence generation process or is identical across the steps of generation. Both methods offer no fine-grained control. To address this concern, we propose Seq-CVAE which learns a latent space for every word position. We encourage this temporal latent space to capture the 'intention' about how to complete the sentence by mimicking a representation which summarizes the future. We illustrate the efficacy of the proposed approach to anticipate the sentence continuation on the challenging MSCOCO dataset, significantly improving diversity metrics compared to baselines while performing on par w.r.t. sentence quality.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext d085b7e8-7f9a-4c53-a2d5-293fe8c26057

Cited by top-tier papers22

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines