MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation
Dawei Wang, Di Zhao, Xinyuan Liu, Marci Chi Ma, Xiaoyang Liu, Chengming Zhou, Gary Ushaw, Richard Davison
Abstract
Credit assignment is a fundamental challenge in cooperative multi-agent reinforcement learning, particularly in embodied AI settings characterized by limited and delayed feedback as well as dynamically changing numbers of active agents. We propose MARS-RA, a framework that reformulates credit assignment as a rank aggregation problem using contributionbased pairwise comparisons among agents generated by large multimodal models. This shift from absolute to relative estimation ensures robustness against noise and dynamic agent participation, converting comparison results into contribution scores for potential-based reward shaping. We provide theoretical justification for the convergence and robustness of the proposed framework, and show that Shapley values can be used as an interpretive reference. Experimental results on challenging tasks of different types indicate that MARS-RA can guide agents toward effective cooperation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2676c09e-76e6-4140-a5ee-6041747d70b7Builds on9
- Shapley Q-Value: A Local Reward Approach to Solve Global Reward GamesJianhong Wang, Yuan Zhang, Tae-Kyun Kim, Yunjie GuAAAI 2020 · 159 citations
- Vision-Language Models are Zero-Shot Reward Models for Reinforcement LearningJuan Rocamonde, Victoriano Montesinos, Elvis Nava, Ethan Perez et al.ICLR 2024 · 154 citations
- SHAQ: Incorporating Shapley Value Theory into Multi-Agent Q-LearningJianhong Wang, Yuan Zhang, Yunjie Gu, Tae-Kyun KimNeurIPS 2022 · 50 citations
- Interleaved Latent Visual Reasoning with Selective Perceptual ModelingShuai Dong, Siyuan Wang, Xingyu Liu, Chenglin Li et al.ACL 2026 · 20 citations
- VideoPro: Adaptive Program Reasoning for Long Video UnderstandingChenglin Li, Feng Han, Yikun Wang, Ruilin Li et al.ACL 2026 · 4 citations
Related papers
- STAS: Spatial-Temporal Return Decomposition for Solving Sparse Rewards Problems in Multi-agent Reinforcement LearningSirui Chen, Zhaowei Zhang, Yaodong Yang, Yali DuAAAI 2024 · 11 citations
- Shapley Counterfactual Credits for Multi-Agent Reinforcement LearningJiahui Li, Kun Kuang, Baoxiang Wang, Furui Liu et al.KDD 2021 · 49 citations
- Shapley-Coop: Credit Assignment for Emergent Cooperation in Self-Interested LLM AgentsYun Hua, Haosheng Chen, Shiqin Wang, Wenhao Li et al.NeurIPS 2025 · 13 citations
- Proactive Multi-Camera Collaboration for 3D Human Pose EstimationHai Ci, Mickel Liu, Xuehai Pan, Fangwei Zhong et al.ICLR 2023 · 6 citations
- Stochastic Self-Organization in Multi-Agent SystemsNurbek Tastan, Samuel Horváth, Karthik NandakumarICLR 2026 · 9 citations
