Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Combining Model-Based and Model-Free Updates for Trajectory-Centric Reinforcement Learning

Jun 18, 2017

Yevgen Chebotar, Karol Hausman, Marvin Zhang, Gaurav Sukhatme, Stefan Schaal, Sergey Levine

Figure 1 for Combining Model-Based and Model-Free Updates for Trajectory-Centric Reinforcement Learning

Figure 2 for Combining Model-Based and Model-Free Updates for Trajectory-Centric Reinforcement Learning

Figure 3 for Combining Model-Based and Model-Free Updates for Trajectory-Centric Reinforcement Learning

Figure 4 for Combining Model-Based and Model-Free Updates for Trajectory-Centric Reinforcement Learning

Share this with someone who'll enjoy it:

Abstract:Reinforcement learning (RL) algorithms for real-world robotic applications need a data-efficient learning process and the ability to handle complex, unknown dynamical systems. These requirements are handled well by model-based and model-free RL approaches, respectively. In this work, we aim to combine the advantages of these two types of methods in a principled manner. By focusing on time-varying linear-Gaussian policies, we enable a model-based algorithm based on the linear quadratic regulator (LQR) that can be integrated into the model-free framework of path integral policy improvement (PI2). We can further combine our method with guided policy search (GPS) to train arbitrary parameterized policies such as deep neural networks. Our simulation and real-world experiments demonstrate that this method can solve challenging manipulation tasks with comparable or better performance than model-free methods while maintaining the sample efficiency of model-based methods. A video presenting our results is available at https://sites.google.com/site/icml17pilqr

* Paper accepted to the International Conference on Machine Learning (ICML) 2017

View paper on

Share this with someone who'll enjoy it:

Title:Combining Model-Based and Model-Free Updates for Trajectory-Centric Reinforcement Learning

Paper and Code