Human-Like Controllable Image Captioning With Verb-Specific Semantic Roles
Long Chen, Zhihong Jiang, Jun Xiao, Wei Liu
2021Year
15Top-tier citations
Abstract
Cap: a man riding a wave on a surfboard. CS: Cap: a man riding a wave on a surfboard in his hand in the sky. CS: level 3 (15-19) Cap: a group of people sitting next to each other in front of a tree. CS: level 4 (20-25) Cap: a group of men and two boys are standing in front of a refrigerator in front of a house.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Generating Visual Spatial Description via Holistic 3D Scene UnderstandingYu Zhao, Hao Fei, Wei Ji, Jianguo Wei et al.ACL 2023 · 40 citations
- DeeCap: Dynamic Early Exiting for Efficient Image CaptioningZhengcong Fei, Xu Yan, Shuhui Wang, Qi TianCVPR 2022 · 39 citations
- Classification-Then-Grounding: Reformulating Video Scene Graphs as Temporal Bipartite GraphsKaifeng Gao, Long Chen, Yulei Niu, Jian Shao et al.CVPR 2022 · 34 citations
- Learning Distinct and Representative Modes for Image CaptioningQi Chen, Chaorui Deng, Qi WuNeurIPS 2022 · 27 citations
- Collaborative Transformers for Grounded Situation RecognitionJunhyeong Cho, Youngseok Yoon, Suha KwakCVPR 2022 · 23 citations
Builds on8
- Learning to Collocate Neural Modules for Image CaptioningXu Yang, Hanwang Zhang, Jianfei CaiICCV 2019 · 84 citations
- Sequential Latent Spaces for Modeling the Intention During Diverse Image CaptioningJyoti Aneja, Harsh Agrawal, Dhruv Batra, Alexander G. SchwingICCV 2019 · 71 citations
- Generating Diverse and Descriptive Image Captions Using Visual ParaphrasesLixin Liu, Jiajun Tang, Xiaojun Wan, Zongming GuoICCV 2019 · 48 citations
- Mixture-Kernel Graph Attention Network for Situation RecognitionMohammed Suhail, Leonid SigalICCV 2019 · 32 citations
- Controllable Video Captioning with an Exemplar SentenceYitian Yuan, Lin Ma, Jingwen Wang, Wenwu ZhuACM MM 2020 · 21 citations
Related papers
- GRES: Generalized Referring Expression SegmentationChang Liu, Henghui Ding, Xudong JiangCVPR 2023
- ArtAdapter: Text-to-Image Style Transfer using Multi-Level Style Encoder and Explicit AdaptationDar-Yen Chen, Hamish Tennent, Ching-Wen HsuCVPR 2024
- ReGenHOI: Unifying Reconstruction and Generation for 3D Human-Object Interaction UnderstandingMiao Xu, Xiangyu Zhu, Zidu Wang, Xusheng Liang et al.CVPR 2026
- CapHuman: Capture Your Moments in Parallel UniversesChao Liang, Fan Ma, Linchao Zhu, Yingying Deng et al.CVPR 2024
- Collaborative Diffusion for Multi-Modal Face Generation and EditingZiqi Huang, Kelvin C. K. Chan, Yuming Jiang, Ziwei LiuCVPR 2023
