Learning Long- and Short-Term User Literal-Preference with Multimodal Hierarchical Transformer Network for Personalized Image Caption
Wei Zhang, Yue Ying, Pan Lu, Hongyuan Zha
Abstract
Personalized image caption, a natural extension of the standard image caption task, requires to generate brief image descriptions tailored for users' writing style and traits, and is more practical to meet users' real demands. Only a few recent studies shed light on this crucial task and learn static user representations to capture their long-term literal-preference. However, it is insufficient to achieve satisfactory performance due to the intrinsic existence of not only long-term user literal-preference, but also short-term literal-preference which is associated with users' recent states. To bridge this gap, we develop a novel multimodal hierarchical transformer network (MHTN) for personalized image caption in this paper. It learns short-term user literal-preference based on users' recent captions through a short-term user encoder at the low level. And at the high level, the multimodal encoder integrates target image representations with short-term literal-preference, as well as long-term literal-preference learned from user IDs. These two encoders enjoy the advantages of the powerful transformer networks. Extensive experiments on two real datasets show the effectiveness of considering two types of user literal-preference simultaneously and better performance over the state-of-the-art models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 96e1a259-322a-4e97-b688-a525416d4513Cited by top-tier papers5
- Personalized Generation In Large Model Era: A SurveyYiyan Xu, Jinghao Zhang, Alireza Salemi, Xinting Hu et al.ACL 2025 · 45 citations
- Betrayed by Captions: Joint Caption Grounding and Generation for Open Vocabulary Instance SegmentationJianzong Wu, Xiangtai Li, Henghui Ding, Xia Li et al.ICCV 2023 · 36 citations
- AESOP: Abstract Encoding of Stories, Objects, and PicturesHareesh Ravi, Kushal Kafle, Scott Cohen, Jonathan Brandt et al.ICCV 2021 · 19 citations
- Personalized Image Descriptions from Attention SequencesRuoyu Xue, Hieu Le, Jingyi Xu, Sounak Mondal et al.CVPR 2026 · 2 citations
- SceneTrilogy: On Human Scene-Sketch and its Complementarity with Photo and TextPinaki Nath Chowdhury, Ayan Kumar Bhunia, Aneeshan Sain, Subhadeep Koley et al.CVPR 2023
Related papers
- Meshed-Memory Transformer for Image CaptioningMarcella Cornia, Matteo Stefanini, Lorenzo Baraldi, Rita CucchiaraCVPR 2020
- Detach and Attach: Stylized Image Captioning without Paired Stylized DatasetYutong Tan, Zheng Lin, Peng Fu, Mingyu Zheng et al.ACM MM 2022 · 8 citations
- Improving Image Captioning through Visual and Semantic Mutual PromotionJing Zhang, Yingshuai Xie, Xiaoqiang LiuACM MM 2023 · 4 citations
- Building User-oriented Personalized Machine Translator based on User-Generated Textual ContentPeng Zhang, Zhengqing Guan, Baoxi Liu, Sharon Xianghua Ding et al.CSCW 2022 · 6 citations
- Multimodal Attention with Image Text Spatial Relationship for OCR-Based Image CaptioningJing Wang, Jinhui Tang, Jiebo LuoACM MM 2020 · 55 citations
