Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Shih-wei Liao

RIIT: Rethinking the Importance of Implementation Tricks in Multi-Agent Reinforcement Learning

Mar 01, 2021

Jian Hu, Haibin Wu, Seth Austin Harding, Siyang Jiang, Shih-wei Liao

Figure 1 for RIIT: Rethinking the Importance of Implementation Tricks in Multi-Agent Reinforcement Learning

Figure 2 for RIIT: Rethinking the Importance of Implementation Tricks in Multi-Agent Reinforcement Learning

Figure 3 for RIIT: Rethinking the Importance of Implementation Tricks in Multi-Agent Reinforcement Learning

Figure 4 for RIIT: Rethinking the Importance of Implementation Tricks in Multi-Agent Reinforcement Learning

Abstract:In recent years, Multi-Agent Deep Reinforcement Learning (MADRL) has been successfully applied to various complex scenarios such as computer games and robot swarms. We investigate the impact of "implementation tricks" of state-of-the-art (SOTA) QMIX-based algorithms. Firstly, we find that such tricks, described as auxiliary details to the core algorithm, seemingly of secondary importance, have a major impact. Our finding demonstrates that, after minimal tuning, QMIX attains extraordinarily high win rates and achieves SOTA in the StarCraft Multi-Agent Challenge (SMAC). Furthermore, we find QMIX's monotonicity condition helps improve sample efficiency in some cooperative tasks, and we propose a new policy-based algorithm, called: RIIT, to prove the importance of the monotonicity condition. RIIT also achieves SOTA in policy-based algorithms. At last, we propose a hypothesis to explain the monotonicity condition. We open-sourced the code at \url{https://github.com/hijkzzz/pymarl2}.

* fix errors

Via

Access Paper or Ask Questions

QR-MIX: Distributional Value Function Factorisation for Cooperative Multi-Agent Reinforcement Learning

Sep 28, 2020

Jian Hu, Seth Austin Harding, Haibin Wu, Shih-wei Liao

Figure 1 for QR-MIX: Distributional Value Function Factorisation for Cooperative Multi-Agent Reinforcement Learning

Figure 2 for QR-MIX: Distributional Value Function Factorisation for Cooperative Multi-Agent Reinforcement Learning

Figure 3 for QR-MIX: Distributional Value Function Factorisation for Cooperative Multi-Agent Reinforcement Learning

Figure 4 for QR-MIX: Distributional Value Function Factorisation for Cooperative Multi-Agent Reinforcement Learning

Abstract:In Cooperative Multi-Agent Reinforcement Learning (MARL) and under the setting of Centralized Training with Decentralized Execution (CTDE), agents observe and interact with their environment locally and independently. With local observation and random sampling, the randomness in rewards and observations leads to randomness in long-term returns. Existing methods such as Value Decomposition Network (VDN) and QMIX estimate the mean value of long-term returns while ignoring randomness. Our proposed model QR-MIX introduces quantile regression, modeling joint state-action values as a distribution, combining QMIX with Implicit Quantile Network (IQN). Besides, because the monotonicity in QMIX limits the expression of joint state-action value distribution and may lead to incorrect estimation results in nonmonotonic cases, we design a flexible loss function to replace the absolute weights found in QMIX. Our methods enhance the expressiveness of our mixing network and are more tolerant of randomness and nonmonotonicity. The experiments demonstrate that QR-MIX outperforms prior works in the StarCraft Multi-Agent Challenge (SMAC) environment.

* There are some important errors in the experiment of this article... I will re-experiment and change the content and title extensively

Via

Access Paper or Ask Questions

An Evaluation of Bitcoin Address Classification based on Transaction History Summarization

Mar 19, 2019

Yu-Jing Lin, Po-Wei Wu, Cheng-Han Hsu, I-Ping Tu, Shih-wei Liao

Figure 1 for An Evaluation of Bitcoin Address Classification based on Transaction History Summarization

Figure 2 for An Evaluation of Bitcoin Address Classification based on Transaction History Summarization

Figure 3 for An Evaluation of Bitcoin Address Classification based on Transaction History Summarization

Figure 4 for An Evaluation of Bitcoin Address Classification based on Transaction History Summarization

Abstract:Bitcoin is a cryptocurrency that features a distributed, decentralized and trustworthy mechanism, which has made Bitcoin a popular global transaction platform. The transaction efficiency among nations and the privacy benefiting from address anonymity of the Bitcoin network have attracted many activities such as payments, investments, gambling, and even money laundering in the past decade. Unfortunately, some criminal behaviors which took advantage of this platform were not identified. This has discouraged many governments to support cryptocurrency. Thus, the capability to identify criminal addresses becomes an important issue in the cryptocurrency network. In this paper, we propose new features in addition to those commonly used in the literature to build a classification model for detecting abnormality of Bitcoin network addresses. These features include various high orders of moments of transaction time (represented by block height) which summarizes the transaction history in an efficient way. The extracted features are trained by supervised machine learning methods on a labeling category data set. The experimental evaluation shows that these features have improved the performance of Bitcoin address classification significantly. We evaluate the results under eight classifiers and achieve the highest Micro-F1/Macro-F1 of 87%/86% with LightGBM.

* 8 pages; accepted by ICBC 2019

Via

Access Paper or Ask Questions