Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Parham Mohammad Panahi

Investigating the Interplay of Prioritized Replay and Generalization

Jul 12, 2024

Parham Mohammad Panahi, Andrew Patterson, Martha White, Adam White

Figure 1 for Investigating the Interplay of Prioritized Replay and Generalization

Figure 2 for Investigating the Interplay of Prioritized Replay and Generalization

Figure 3 for Investigating the Interplay of Prioritized Replay and Generalization

Figure 4 for Investigating the Interplay of Prioritized Replay and Generalization

Abstract:Experience replay is ubiquitous in reinforcement learning, to reuse past data and improve sample efficiency. Though a variety of smart sampling schemes have been introduced to improve performance, uniform sampling by far remains the most common approach. One exception is Prioritized Experience Replay (PER), where sampling is done proportionally to TD errors, inspired by the success of prioritized sweeping in dynamic programming. The original work on PER showed improvements in Atari, but follow-up results are mixed. In this paper, we investigate several variations on PER, to attempt to understand where and when PER may be useful. Our findings in prediction tasks reveal that while PER can improve value propagation in tabular settings, behavior is significantly different when combined with neural networks. Certain mitigations -- like delaying target network updates to control generalization and using estimates of expected TD errors in PER to avoid chasing stochasticity -- can avoid large spikes in error with PER and neural networks, but nonetheless generally do not outperform uniform replay. In control tasks, none of the prioritized variants consistently outperform uniform replay.

* Published in the Reinforcement Learning Conference 2024

Via

Access Paper or Ask Questions

A New View on Planning in Online Reinforcement Learning

Jun 03, 2024

Kevin Roice, Parham Mohammad Panahi, Scott M. Jordan, Adam White, Martha White

Figure 1 for A New View on Planning in Online Reinforcement Learning

Figure 2 for A New View on Planning in Online Reinforcement Learning

Figure 3 for A New View on Planning in Online Reinforcement Learning

Figure 4 for A New View on Planning in Online Reinforcement Learning

Abstract:This paper investigates a new approach to model-based reinforcement learning using background planning: mixing (approximate) dynamic programming updates and model-free updates, similar to the Dyna architecture. Background planning with learned models is often worse than model-free alternatives, such as Double DQN, even though the former uses significantly more memory and computation. The fundamental problem is that learned models can be inaccurate and often generate invalid states, especially when iterated many steps. In this paper, we avoid this limitation by constraining background planning to a set of (abstract) subgoals and learning only local, subgoal-conditioned models. This goal-space planning (GSP) approach is more computationally efficient, naturally incorporates temporal abstraction for faster long-horizon planning and avoids learning the transition dynamics entirely. We show that our GSP algorithm can propagate value from an abstract space in a manner that helps a variety of base learners learn significantly faster in different domains.

* Published in the Planning and Reinforcement Learning Workshop at ICAPS 2024. arXiv admin note: text overlap with arXiv:2206.02902

Via

Access Paper or Ask Questions

Tuning for the Unknown: Revisiting Evaluation Strategies for Lifelong RL

Apr 02, 2024

Golnaz Mesbahi, Olya Mastikhina, Parham Mohammad Panahi, Martha White, Adam White

Figure 1 for Tuning for the Unknown: Revisiting Evaluation Strategies for Lifelong RL

Figure 2 for Tuning for the Unknown: Revisiting Evaluation Strategies for Lifelong RL

Figure 3 for Tuning for the Unknown: Revisiting Evaluation Strategies for Lifelong RL

Figure 4 for Tuning for the Unknown: Revisiting Evaluation Strategies for Lifelong RL

Abstract:In continual or lifelong reinforcement learning access to the environment should be limited. If we aspire to design algorithms that can run for long-periods of time, continually adapting to new, unexpected situations then we must be willing to deploy our agents without tuning their hyperparameters over the agent's entire lifetime. The standard practice in deep RL -- and even continual RL -- is to assume unfettered access to deployment environment for the full lifetime of the agent. This paper explores the notion that progress in lifelong RL research has been held back by inappropriate empirical methodologies. In this paper we propose a new approach for tuning and evaluating lifelong RL agents where only one percent of the experiment data can be used for hyperparameter tuning. We then conduct an empirical study of DQN and Soft Actor Critic across a variety of continuing and non-stationary domains. We find both methods generally perform poorly when restricted to one-percent tuning, whereas several algorithmic mitigations designed to maintain network plasticity perform surprising well. In addition, we find that properties designed to measure the network's ability to learn continually indeed correlate with performance under one-percent tuning.

Via

Access Paper or Ask Questions