VrR-VG: Refocusing Visually-Relevant Relationships
Yuanzhi Liang, Yalong Bai, Wei Zhang, Xueming Qian, Li Zhu, Tao Mei
摘要
Relationships encode the interactions among individual instances and play a critical role in deep visual scene understanding. Suffering from the high predictability with non-visual information, relationship models tend to fit the statistical bias rather than ``learning" to infer the relationships from images. To encourage further development in visual relationships, we propose a novel method to mine more valuable relationships by automatically pruning visually-irrelevant relationships. We construct a new scene graph dataset named Visually-Relevant Relationships Dataset (VrR-VG) based on Visual Genome. Compared with existing datasets, the performance gap between learnable and statistical method is more significant in VrR-VG, and frequency-based analysis does not work anymore. Moreover, we propose to learn a relationship-aware representation by jointly considering instances, attributes and relationships. By applying the representation-aware feature learned on VrR-VG, the performances of image captioning and visual question answering are systematically improved, which demonstrates the effectiveness of both our dataset and features embedding schema. Both our VrR-VG dataset and representation-aware features will be made publicly available soon.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- From General to Specific: Informative Scene Graph Generation via Balance AdjustmentYuyu Guo, Lianli Gao, Xuanhan Wang, Yuxuan Hu 等ICCV 2021 · 被引用 96 次
- Image-to-Image Retrieval by Learning Similarity between Scene GraphsSangwoong Yoon, Woo-Young Kang, Sungwook Jeon, SeongEun Lee 等AAAI 2021 · 被引用 57 次
- Fine-Grained Predicates Learning for Scene Graph GenerationXinyu Lyu, Lianli Gao, Yuyu Guo, Zhou Zhao 等CVPR 2022 · 被引用 48 次
- Instruction-Guided Visual MaskingJinliang Zheng, Jianxiong Li, Sijie Cheng, Yinan Zheng 等NeurIPS 2024 · 被引用 22 次
- Topic Scene Graph Generation by Attention Distillation from CaptionWenbin Wang, Ruiping Wang, Xilin ChenICCV 2021 · 被引用 16 次
相关 Paper
- AI-VQA: Visual Question Answering based on Agent Interaction with InterpretabilityRengang Li, Cong Xu, Zhenhua Guo, Baoyu Fan 等ACM MM 2022 · 被引用 7 次
- Scene Graph Prediction With Limited LabelsRanjay Krishna, Vincent S. Chen, Paroma Varma, Michael S. Bernstein 等ICCV 2019 · 被引用 5 次
- Memory-Based Network for Scene Graph with Unbalanced RelationsWeitao Wang, Ruyang Liu, Meng Wang, Sen Wang 等ACM MM 2020 · 被引用 11 次
- Not All Relations are Equal: Mining Informative Labels for Scene Graph GenerationArushi Goel, Basura Fernando, Frank Keller, Hakan BilenCVPR 2022 · 被引用 30 次
- Inferential Visual Question GenerationChao Bi, Shuhui Wang, Zhe Xue, Shengbo Chen 等ACM MM 2022 · 被引用 8 次
