Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Scott Jordan

Soft Options Critic

Jun 11, 2019

Elita Lobo, Scott Jordan

Abstract:The option-critic architecture (Bacon, Harb, and Precup 2017) and several variants have successfully demonstrated the use of the options framework proposed by Sutton et al (Sutton, Precup, and Singh1999) to scale learning and planning in hierarchical tasks. Although most of these frameworks use entropy as a regularizer to improve exploration, they do not maximize entropy along with returns at every time step. (Haarnoja et al., 2018d) recently introduced an off-policy actor critic algorithm in theSoft Actor Critic paper that maximize returns while maximizing entropy in a constrained manner thus enabling learning of robust options in continuous and discrete action spaces In this paper we adopt the architecture of soft-actor critic to investigate the effect of maximizing entropy of each options and inter-option policy in options framework. We derive the soft options improvement theorem and propose a novel soft-options framework to incorporate maximization of entropy of actions and options in a constrained manner. Our experiments show that the modified options-critic framework generates robust policies which allows fast recovery when environment is subjected to perturbations and outperforms vanilla options-critic framework in most hierarchical tasks

* In the current version of the paper, there is an error in the definition of the value function, unintended text overlap in the environment description, and incomplete experimentation. These changes will take a significant amount of time to address. Thus, we are withdrawing the paper until these changes can be implemented

Via

Access Paper or Ask Questions

Learning Action Representations for Reinforcement Learning

Feb 01, 2019

Yash Chandak, Georgios Theocharous, James Kostas, Scott Jordan, Philip S. Thomas

Figure 1 for Learning Action Representations for Reinforcement Learning

Figure 2 for Learning Action Representations for Reinforcement Learning

Figure 3 for Learning Action Representations for Reinforcement Learning

Figure 4 for Learning Action Representations for Reinforcement Learning

Abstract:Most model-free reinforcement learning methods leverage state representations (embeddings) for generalization, but either ignore structure in the space of actions or assume the structure is provided a priori. We show how a policy can be decomposed into a component that acts in a low-dimensional space of action representations and a component that transforms these representations into actual actions. These representations improve generalization over large, finite action sets by allowing the agent to infer the outcomes of actions similar to actions already taken. We provide an algorithm to both learn and use action representations and provide conditions for its convergence. The efficacy of the proposed method is demonstrated on large-scale real-world problems.

Via

Access Paper or Ask Questions