Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Training Discrete Deep Generative Models via Gapped Straight-Through Estimator

Jun 15, 2022

Ting-Han Fan, Ta-Chung Chi, Alexander I. Rudnicky, Peter J. Ramadge

Figure 1 for Training Discrete Deep Generative Models via Gapped Straight-Through Estimator

Figure 2 for Training Discrete Deep Generative Models via Gapped Straight-Through Estimator

Figure 3 for Training Discrete Deep Generative Models via Gapped Straight-Through Estimator

Figure 4 for Training Discrete Deep Generative Models via Gapped Straight-Through Estimator

Share this with someone who'll enjoy it:

Abstract:While deep generative models have succeeded in image processing, natural language processing, and reinforcement learning, training that involves discrete random variables remains challenging due to the high variance of its gradient estimation process. Monte Carlo is a common solution used in most variance reduction approaches. However, this involves time-consuming resampling and multiple function evaluations. We propose a Gapped Straight-Through (GST) estimator to reduce the variance without incurring resampling overhead. This estimator is inspired by the essential properties of Straight-Through Gumbel-Softmax. We determine these properties and show via an ablation study that they are essential. Experiments demonstrate that the proposed GST estimator enjoys better performance compared to strong baselines on two discrete deep generative modeling tasks, MNIST-VAE and ListOps.

* Accepted at the International Conference on Machine Learning (ICML) 2022. The first two authors contributed equally

View paper on

Share this with someone who'll enjoy it:

Title:Training Discrete Deep Generative Models via Gapped Straight-Through Estimator

Paper and Code