Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

James Kostas

Classical Policy Gradient: Preserving Bellman's Principle of Optimality

Jun 06, 2019

Philip S. Thomas, Scott M. Jordan, Yash Chandak, Chris Nota, James Kostas

Abstract:We propose a new objective function for finite-horizon episodic Markov decision processes that better captures Bellman's principle of optimality, and provide an expression for the gradient of the objective.

* 1 page, 0 figures

Via

Access Paper or Ask Questions

Asynchronous Coagent Networks: Stochastic Networks for Reinforcement Learning without Backpropagation or a Clock

Feb 21, 2019

James Kostas, Chris Nota, Philip S. Thomas

Figure 1 for Asynchronous Coagent Networks: Stochastic Networks for Reinforcement Learning without Backpropagation or a Clock

Figure 2 for Asynchronous Coagent Networks: Stochastic Networks for Reinforcement Learning without Backpropagation or a Clock

Figure 3 for Asynchronous Coagent Networks: Stochastic Networks for Reinforcement Learning without Backpropagation or a Clock

Figure 4 for Asynchronous Coagent Networks: Stochastic Networks for Reinforcement Learning without Backpropagation or a Clock

Abstract:In this paper we introduce a reinforcement learning (RL) approach for training policies, including artificial neural network policies, that is both backpropagation-free and clock-free. It is backpropagation-free in that it does not propagate any information backwards through the network. It is clock-free in that no signal is given to each node in the network to specify when it should compute its output and when it should update its weights. We contend that these two properties increase the biological plausibility of our algorithms and facilitate distributed implementations. Additionally, our approach eliminates the need for customized learning rules for hierarchical RL algorithms like the option-critic.

* Removed LaTeX commands from metadata abstract. Corrected typo in section 4. Changed Title

Via

Access Paper or Ask Questions

Learning Action Representations for Reinforcement Learning

Feb 01, 2019

Yash Chandak, Georgios Theocharous, James Kostas, Scott Jordan, Philip S. Thomas

Figure 1 for Learning Action Representations for Reinforcement Learning

Figure 2 for Learning Action Representations for Reinforcement Learning

Figure 3 for Learning Action Representations for Reinforcement Learning

Figure 4 for Learning Action Representations for Reinforcement Learning

Abstract:Most model-free reinforcement learning methods leverage state representations (embeddings) for generalization, but either ignore structure in the space of actions or assume the structure is provided a priori. We show how a policy can be decomposed into a component that acts in a low-dimensional space of action representations and a component that transforms these representations into actual actions. These representations improve generalization over large, finite action sets by allowing the agent to infer the outcomes of actions similar to actions already taken. We provide an algorithm to both learn and use action representations and provide conditions for its convergence. The efficacy of the proposed method is demonstrated on large-scale real-world problems.

Via

Access Paper or Ask Questions