Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Jiancong Huang

Hyperparameter Auto-tuning in Self-Supervised Robotic Learning

Oct 19, 2020

Jiancong Huang, Juan Rojas, Matthieu Zimmer, Hongmin Wu, Yisheng Guan, Paul Weng

Figure 1 for Hyperparameter Auto-tuning in Self-Supervised Robotic Learning

Figure 2 for Hyperparameter Auto-tuning in Self-Supervised Robotic Learning

Figure 3 for Hyperparameter Auto-tuning in Self-Supervised Robotic Learning

Figure 4 for Hyperparameter Auto-tuning in Self-Supervised Robotic Learning

Abstract:Policy optimization in reinforcement learning requires the selection of numerous hyperparameters across different environments. Fixing them incorrectly may negatively impact optimization performance leading notably to insufficient or redundant learning. Insufficient learning (due to convergence to local optima) results in under-performing policies whilst redundant learning wastes time and resources. The effects are further exacerbated when using single policies to solve multi-task learning problems. In this paper, we study how the Evidence Lower Bound (ELBO) used in Variational Auto-Encoders (VAEs) is affected by the diversity of image samples. Different tasks or setups in visual reinforcement learning incur varying diversity. We exploit the ELBO to create an auto-tuning technique in self-supervised reinforcement learning. Our approach can auto-tune three hyperparameters: the replay buffer size, the number of policy gradient updates during each epoch, and the number of exploration steps during each epoch. We use the state-of-the-art self-supervised robotic learning framework (Reinforcement Learning with Imagined Goals (RIG) using Soft Actor-Critic) as baseline for experimental verification. Experiments show that our method can auto-tune online and yields the best performance at a fraction of the time and computational resources. Code, video, and appendix for simulated and real-robot experiments can be found at http://www.JuanRojas.net/autotune.

Via

Access Paper or Ask Questions

Towards More Sample Efficiency in Reinforcement Learning with Data Augmentation

Nov 15, 2019

Yijiong Lin, Jiancong Huang, Matthieu Zimmer, Juan Rojas, Paul Weng

Figure 1 for Towards More Sample Efficiency in Reinforcement Learning with Data Augmentation

Figure 2 for Towards More Sample Efficiency in Reinforcement Learning with Data Augmentation

Figure 3 for Towards More Sample Efficiency in Reinforcement Learning with Data Augmentation

Abstract:Deep reinforcement learning (DRL) is a promising approach for adaptive robot control, but its current application to robotics is currently hindered by high sample requirements. We propose two novel data augmentation techniques for DRL in order to reuse more efficiently observed data. The first one called Kaleidoscope Experience Replay exploits reflectional symmetries, while the second called Goal-augmented Experience Replay takes advantage of lax goal definitions. Our preliminary experimental results show a large increase in learning speed.

* NeurIPS 2019 Workshop on Robot Learning: Control and Interaction in the Real World (accepted after double-blind peer review). arXiv admin note: substantial text overlap with arXiv:1909.10707

Via

Access Paper or Ask Questions

Invariant Transform Experience Replay

Oct 10, 2019

Yijiong Lin, Jiancong Huang, Matthieu Zimmer, Juan Rojas, Paul Weng

Figure 1 for Invariant Transform Experience Replay

Figure 2 for Invariant Transform Experience Replay

Figure 3 for Invariant Transform Experience Replay

Figure 4 for Invariant Transform Experience Replay

Abstract:Deep reinforcement learning (DRL) is a promising approach for adaptive robot control, but its current application to robotics is currently hindered by high sample requirements. We propose two novel data augmentation techniques for DRL based on invariant transformations of trajectories in order to reuse more efficiently observed interaction. The first one called Kaleidoscope Experience Replay exploits reflectional symmetries, while the second called Goal-augmented Experience Replay takes advantage of lax goal definitions. In the Fetch tasks from OpenAI Gym, our experimental results show a large increase in learning speed.

* 7 pages, 6 figures

Via

Access Paper or Ask Questions