Federated Learning for Vision-and-Language Grounding Problems
Fenglin Liu, Xian Wu, Shen Ge, Wei Fan, Yuexian Zou
Abstract
Recently, vision-and-language grounding problems, e.g., image captioning and visual question answering (VQA), has attracted extensive interests from both academic and industrial worlds. However, given the similarity of these tasks, the efforts to obtain better results by combining the merits of their algorithms are not well studied. Inspired by the recent success of federated learning, we propose a federated learning framework to obtain various types of image representations from different tasks, which are then fused together to form fine-grained image representations. The representations merge useful features from different vision-and-language grounding problems, and are thus much more powerful than the original representations alone in individual tasks. To learn such image representations, we propose the Aligning, Integrating and Mapping Network (aimNet). The aimNet is validated on three federated learning settings, which include horizontal federated learning, vertical federated learning, and federated transfer learning. Experiments of aimNet-based federated learning framework on two representative tasks, i.e., image captioning and VQA, demonstrate the effective and universal improvements of all metrics over the baselines. In image captioning, we are able to get 14% and 13% relative gain on the task-specific metrics CIDEr and SPICE, respectively. In VQA, we could also boost the performance of strong baselines by up to 3%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 075ad978-2a34-489a-b32e-4e258f2fe980Cited by top-tier papers17
- Provably Secure Federated Learning against Malicious ClientsXiaoyu Cao, Jinyuan Jia, Neil Zhenqiang GongAAAI 2021 · 161 citations
- Tailor Versatile Multi-Modal Learning for Multi-Label Emotion RecognitionYi Zhang, Mingyuan Chen, Jundong Shen, Chongjun WangAAAI 2022 · 92 citations
- Towards Optimal Multi-Modal Federated Learning on Non-IID Data with Hierarchical Gradient BlendingSijia Chen, Baochun LiINFOCOM 2022 · 57 citations
- Prophet Attention: Predicting Attention with Future AttentionFenglin Liu, Xuancheng Ren, Xian Wu, Shen Ge et al.NeurIPS 2020 · 52 citations
- FairVFL: A Fair Vertical Federated Learning Framework with Contrastive Adversarial LearningTao Qi, Fangzhao Wu, Chuhan Wu, Lingjuan Lyu et al.NeurIPS 2022 · 51 citations
Builds on1
Related papers
- 12-in-1: Multi-Task Vision and Language Representation LearningJiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh et al.CVPR 2020
- 3DJCG: A Unified Framework for Joint Dense Captioning and Visual Grounding on 3D Point CloudsDaigang Cai, Lichen Zhao, Jing Zhang, Lu Sheng et al.CVPR 2022 · 80 citations
- Remodeling Semantic Relationships in Vision-Language Fine-TuningXiangyang Wu, Liu Liu, Baosheng Yu, Jiayan Qiu et al.AAAI 2026
- Unifying Vision-and-Language Tasks via Text GenerationJaemin Cho, Jie Lei, Hao Tan, Mohit BansalICML 2021 · 624 citations
- Language Features Matter: Effective Language Representations for Vision-Language TasksAndrea Burns, Reuben Tan, Kate Saenko, Stan Sclaroff et al.ICCV 2019 · 28 citations
