Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning
Jyoti Aneja, Harsh Agrawal, Dhruv Batra, Alexander G. Schwing
Abstract
Diverse and accurate vision+language modeling is an important goal to retain creative freedom and maintain user engagement. However, adequately capturing the intricacies of diversity in language models is challenging. Recent works commonly resort to latent variable models augmented with more or less supervision from object detectors or part-of-speech tags [10, 40] . Common to all those methods is the fact that the latent variable either only initializes the sentence generation process or is identical across the steps of generation. Both methods offer no fine-grained control. To address this concern, we propose Seq-CVAE which learns a latent space for every word position. We encourage this temporal latent space to capture the 'intention' about how to complete the sentence by mimicking a representation which summarizes the future. We illustrate the efficacy of the proposed approach to anticipate the sentence continuation on the challenging MSCOCO dataset, significantly improving diversity metrics compared to baselines while performing on par w.r.t. sentence quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d085b7e8-7f9a-4c53-a2d5-293fe8c26057Cited by top-tier papers22
- Sequential Latent Knowledge Selection for Knowledge-Grounded DialogueByeongchang Kim, Jaewoo Ahn, Gunhee KimICLR 2020 · 179 citations
- D2C: Diffusion-Decoding Models for Few-Shot Conditional GenerationAbhishek Sinha, Jiaming Song, Chenlin Meng, Stefano ErmonNeurIPS 2021 · 149 citations
- A Contrastive Learning Approach for Training Variational Autoencoder PriorsJyoti Aneja, Alexander G. Schwing, Jan Kautz, Arash VahdatNeurIPS 2021 · 112 citations
- Removing Bias in Multi-modal Classifiers: Regularization by Maximizing Functional EntropiesItai Gat, Idan Schwartz, Alexander G. Schwing, Tamir HazanNeurIPS 2020 · 111 citations
- Diverse Image Captioning with Context-Object Split Latent SpacesShweta Mahajan, Stefan RothNeurIPS 2020 · 47 citations
Related papers
- Contextually Plausible and Diverse 3D Human Motion PredictionSadegh Aliakbarian, Fatemeh Sadat Saleh, Lars Petersson, Stephen Gould et al.ICCV 2021 · 44 citations
- Top-Down Semantic Refinement for Image CaptioningJusheng Zhang, Kaitong Cai, Jing Yang, Jian Wang et al.AAAI 2026 · 16 citations
- Discriminative Latent Semantic Graph for Video CaptioningYang Bai, Junyan Wang, Yang Long, Bingzhang Hu et al.ACM MM 2021 · 26 citations
- Progress-Aware Video Frame CaptioningZihui Xue, Joungbin An, Xitong Yang, Kristen GraumanCVPR 2025
- Chain of World: World Model Thinking in Latent MotionFuxiang Yang, Donglin Di, Lulu Tang, Xuancheng Zhang et al.CVPR 2026 · 11 citations
