Show, Edit and Tell: A Framework for Editing Image Captions
Fawaz Sammani, Luke Melas-Kyriazi
摘要
Most image captioning frameworks generate captions directly from images, learning a mapping from visual features to natural language. However, editing existing captions can be easier than generating new ones from scratch. Intuitively, when editing captions, a model is not required to learn information that is already present in the caption (i.e. sentence structure), enabling it to focus on fixing details (e.g. replacing repetitive words). This paper proposes a novel approach to image captioning based on iterative adaptive refinement of an existing caption. Specifically, our caption-editing model consisting of two sub-modules: (1) EditNet, a language module with an adaptive copy mechanism (Copy-LSTM) and a Selective Copy Memory Attention mechanism (SCMA), and (2) DCNet, an LSTM-based denoising auto-encoder. These components enable our model to directly copy from and modify existing captions. Experiments demonstrate that our new approach achieves stateof-art performance on the MS COCO dataset both with and without sequence-level training. Code can be found at https://github.com/fawazsammani/show-edit-tell .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Non-Autoregressive Coarse-to-Fine Video CaptioningBang Yang, Yuexian Zou, Fenglin Liu, Can ZhangAAAI 2021 · 被引用 92 次
- ReFormer: The Relational Transformer for Image CaptioningXuewen Yang, Yingru Liu, Xin WangACM MM 2022 · 被引用 70 次
- ImageExplorer: Multi-Layered Touch Exploration to Encourage Skepticism Towards Imperfect AI-Generated Image CaptionsJaewook Lee, Jaylin Herskovitz, Yi-Hao Peng, Anhong GuoCHI 2022 · 被引用 55 次
- NLX-GPT: A Model for Natural Language Explanations in Vision and Vision-Language TasksFawaz Sammani, Tanmoy Mukherjee, Nikos DeligiannisCVPR 2022 · 被引用 46 次
- Efficient Modeling of Future Context for Image CaptioningZhengcong FeiACM MM 2022 · 被引用 10 次
它引用的顶会 Paper1
相关 Paper
- Exploring Overall Contextual Information for Image Captioning in Human-Like Cognitive StyleHongwei Ge, Zehang Yan, Kai Zhang, Mingde Zhao 等ICCV 2019 · 被引用 25 次
- Semi-Autoregressive Image CaptioningXu Yan, Zhengcong Fei, Zekang Li, Shuhui Wang 等ACM MM 2021 · 被引用 22 次
- Iterative Back Modification for Faster Image CaptioningZhengcong FeiACM MM 2020 · 被引用 27 次
- Reflective Decoding Network for Image CaptioningLei Ke, Wenjie Pei, Ruiyu Li, Xiaoyong Shen 等ICCV 2019 · 被引用 107 次
- DIFNet: Boosting Visual Information Flow for Image CaptioningMingrui Wu, Xuying Zhang, Xiaoshuai Sun, Yiyi Zhou 等CVPR 2022 · 被引用 68 次
