Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Seth Austin Harding

RIIT: Rethinking the Importance of Implementation Tricks in Multi-Agent Reinforcement Learning

Mar 01, 2021

Jian Hu, Haibin Wu, Seth Austin Harding, Siyang Jiang, Shih-wei Liao

Figure 1 for RIIT: Rethinking the Importance of Implementation Tricks in Multi-Agent Reinforcement Learning

Figure 2 for RIIT: Rethinking the Importance of Implementation Tricks in Multi-Agent Reinforcement Learning

Figure 3 for RIIT: Rethinking the Importance of Implementation Tricks in Multi-Agent Reinforcement Learning

Figure 4 for RIIT: Rethinking the Importance of Implementation Tricks in Multi-Agent Reinforcement Learning

Abstract:In recent years, Multi-Agent Deep Reinforcement Learning (MADRL) has been successfully applied to various complex scenarios such as computer games and robot swarms. We investigate the impact of "implementation tricks" of state-of-the-art (SOTA) QMIX-based algorithms. Firstly, we find that such tricks, described as auxiliary details to the core algorithm, seemingly of secondary importance, have a major impact. Our finding demonstrates that, after minimal tuning, QMIX attains extraordinarily high win rates and achieves SOTA in the StarCraft Multi-Agent Challenge (SMAC). Furthermore, we find QMIX's monotonicity condition helps improve sample efficiency in some cooperative tasks, and we propose a new policy-based algorithm, called: RIIT, to prove the importance of the monotonicity condition. RIIT also achieves SOTA in policy-based algorithms. At last, we propose a hypothesis to explain the monotonicity condition. We open-sourced the code at \url{https://github.com/hijkzzz/pymarl2}.

* fix errors

Via

Access Paper or Ask Questions

QR-MIX: Distributional Value Function Factorisation for Cooperative Multi-Agent Reinforcement Learning

Sep 28, 2020

Jian Hu, Seth Austin Harding, Haibin Wu, Shih-wei Liao

Figure 1 for QR-MIX: Distributional Value Function Factorisation for Cooperative Multi-Agent Reinforcement Learning

Figure 2 for QR-MIX: Distributional Value Function Factorisation for Cooperative Multi-Agent Reinforcement Learning

Figure 3 for QR-MIX: Distributional Value Function Factorisation for Cooperative Multi-Agent Reinforcement Learning

Figure 4 for QR-MIX: Distributional Value Function Factorisation for Cooperative Multi-Agent Reinforcement Learning

Abstract:In Cooperative Multi-Agent Reinforcement Learning (MARL) and under the setting of Centralized Training with Decentralized Execution (CTDE), agents observe and interact with their environment locally and independently. With local observation and random sampling, the randomness in rewards and observations leads to randomness in long-term returns. Existing methods such as Value Decomposition Network (VDN) and QMIX estimate the mean value of long-term returns while ignoring randomness. Our proposed model QR-MIX introduces quantile regression, modeling joint state-action values as a distribution, combining QMIX with Implicit Quantile Network (IQN). Besides, because the monotonicity in QMIX limits the expression of joint state-action value distribution and may lead to incorrect estimation results in nonmonotonic cases, we design a flexible loss function to replace the absolute weights found in QMIX. Our methods enhance the expressiveness of our mixing network and are more tolerant of randomness and nonmonotonicity. The experiments demonstrate that QR-MIX outperforms prior works in the StarCraft Multi-Agent Challenge (SMAC) environment.

* There are some important errors in the experiment of this article... I will re-experiment and change the content and title extensively

Via

Access Paper or Ask Questions