Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Ariyan Bighashdel

Policy Space Response Oracles: A Survey

Mar 04, 2024

Ariyan Bighashdel, Yongzhao Wang, Stephen McAleer, Rahul Savani, Frans A. Oliehoek

Figure 1 for Policy Space Response Oracles: A Survey

Figure 2 for Policy Space Response Oracles: A Survey

Abstract:In game theory, a game refers to a model of interaction among rational decision-makers or players, making choices with the goal of achieving their individual objectives. Understanding their behavior in games is often referred to as game reasoning. This survey provides a comprehensive overview of a fast-developing game-reasoning framework for large games, known as Policy Space Response Oracles (PSRO). We first motivate PSRO, provide historical context, and position PSRO within game-reasoning approaches. We then focus on the strategy exploration issue for PSRO, the challenge of assembling an effective strategy portfolio for modeling the underlying game with minimum computational cost. We also survey current research directions for enhancing the efficiency of PSRO, and explore the applications of PSRO across various domains. We conclude by discussing open questions and future research.

* Ariyan Bighashdel and Yongzhao Wang contributed equally

Via

Access Paper or Ask Questions

Off-Policy Action Anticipation in Multi-Agent Reinforcement Learning

Apr 04, 2023

Ariyan Bighashdel, Daan de Geus, Pavol Jancura, Gijs Dubbelman

Figure 1 for Off-Policy Action Anticipation in Multi-Agent Reinforcement Learning

Figure 2 for Off-Policy Action Anticipation in Multi-Agent Reinforcement Learning

Figure 3 for Off-Policy Action Anticipation in Multi-Agent Reinforcement Learning

Figure 4 for Off-Policy Action Anticipation in Multi-Agent Reinforcement Learning

Abstract:Learning anticipation in Multi-Agent Reinforcement Learning (MARL) is a reasoning paradigm where agents anticipate the learning steps of other agents to improve cooperation among themselves. As MARL uses gradient-based optimization, learning anticipation requires using Higher-Order Gradients (HOG), with so-called HOG methods. Existing HOG methods are based on policy parameter anticipation, i.e., agents anticipate the changes in policy parameters of other agents. Currently, however, these existing HOG methods have only been applied to differentiable games or games with small state spaces. In this work, we demonstrate that in the case of non-differentiable games with large state spaces, existing HOG methods do not perform well and are inefficient due to their inherent limitations related to policy parameter anticipation and multiple sampling stages. To overcome these problems, we propose Off-Policy Action Anticipation (OffPA2), a novel framework that approaches learning anticipation through action anticipation, i.e., agents anticipate the changes in actions of other agents, via off-policy sampling. We theoretically analyze our proposed OffPA2 and employ it to develop multiple HOG methods that are applicable to non-differentiable games with large state spaces. We conduct a large set of experiments and illustrate that our proposed HOG methods outperform the existing ones regarding efficiency and performance.

Via

Access Paper or Ask Questions

Deep Adaptive Multi-Intention Inverse Reinforcement Learning

Jul 14, 2021

Ariyan Bighashdel, Panagiotis Meletis, Pavol Jancura, Gijs Dubbelman

Figure 1 for Deep Adaptive Multi-Intention Inverse Reinforcement Learning

Figure 2 for Deep Adaptive Multi-Intention Inverse Reinforcement Learning

Figure 3 for Deep Adaptive Multi-Intention Inverse Reinforcement Learning

Figure 4 for Deep Adaptive Multi-Intention Inverse Reinforcement Learning

Abstract:This paper presents a deep Inverse Reinforcement Learning (IRL) framework that can learn an a priori unknown number of nonlinear reward functions from unlabeled experts' demonstrations. For this purpose, we employ the tools from Dirichlet processes and propose an adaptive approach to simultaneously account for both complex and unknown number of reward functions. Using the conditional maximum entropy principle, we model the experts' multi-intention behaviors as a mixture of latent intention distributions and derive two algorithms to estimate the parameters of the deep reward network along with the number of experts' intentions from unlabeled demonstrations. The proposed algorithms are evaluated on three benchmarks, two of which have been specifically extended in this study for multi-intention IRL, and compared with well-known baselines. We demonstrate through several experiments the advantages of our algorithms over the existing approaches and the benefits of online inferring, rather than fixing beforehand, the number of expert's intentions.

* Accepted for presentation at ECML/PKDD 2021

Via

Access Paper or Ask Questions