Federated Learning for Vision-and-Language Grounding Problems
Fenglin Liu, Xian Wu, Shen Ge, Wei Fan, Yuexian Zou
摘要
Recently, vision-and-language grounding problems, e.g., image captioning and visual question answering (VQA), has attracted extensive interests from both academic and industrial worlds. However, given the similarity of these tasks, the efforts to obtain better results by combining the merits of their algorithms are not well studied. Inspired by the recent success of federated learning, we propose a federated learning framework to obtain various types of image representations from different tasks, which are then fused together to form fine-grained image representations. The representations merge useful features from different vision-and-language grounding problems, and are thus much more powerful than the original representations alone in individual tasks. To learn such image representations, we propose the Aligning, Integrating and Mapping Network (aimNet). The aimNet is validated on three federated learning settings, which include horizontal federated learning, vertical federated learning, and federated transfer learning. Experiments of aimNet-based federated learning framework on two representative tasks, i.e., image captioning and VQA, demonstrate the effective and universal improvements of all metrics over the baselines. In image captioning, we are able to get 14% and 13% relative gain on the task-specific metrics CIDEr and SPICE, respectively. In VQA, we could also boost the performance of strong baselines by up to 3%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Provably Secure Federated Learning against Malicious ClientsXiaoyu Cao, Jinyuan Jia, Neil Zhenqiang GongAAAI 2021 · 被引用 161 次
- Tailor Versatile Multi-Modal Learning for Multi-Label Emotion RecognitionYi Zhang, Mingyuan Chen, Jundong Shen, Chongjun WangAAAI 2022 · 被引用 92 次
- Towards Optimal Multi-Modal Federated Learning on Non-IID Data with Hierarchical Gradient BlendingSijia Chen, Baochun LiINFOCOM 2022 · 被引用 57 次
- Prophet Attention: Predicting Attention with Future AttentionFenglin Liu, Xuancheng Ren, Xian Wu, Shen Ge 等NeurIPS 2020 · 被引用 52 次
- FairVFL: A Fair Vertical Federated Learning Framework with Contrastive Adversarial LearningTao Qi, Fangzhao Wu, Chuhan Wu, Lingjuan Lyu 等NeurIPS 2022 · 被引用 51 次
它引用的顶会 Paper1
相关 Paper
- 12-in-1: Multi-Task Vision and Language Representation LearningJiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh 等CVPR 2020
- 3DJCG: A Unified Framework for Joint Dense Captioning and Visual Grounding on 3D Point CloudsDaigang Cai, Lichen Zhao, Jing Zhang, Lu Sheng 等CVPR 2022 · 被引用 80 次
- Remodeling Semantic Relationships in Vision-Language Fine-TuningXiangyang Wu, Liu Liu, Baosheng Yu, Jiayan Qiu 等AAAI 2026
- Unifying Vision-and-Language Tasks via Text GenerationJaemin Cho, Jie Lei, Hao Tan, Mohit BansalICML 2021 · 被引用 624 次
- Language Features Matter: Effective Language Representations for Vision-Language TasksAndrea Burns, Reuben Tan, Kate Saenko, Stan Sclaroff 等ICCV 2019 · 被引用 28 次
