Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Vlad Zhukov

Differentiable lower bound for expected BLEU score

Aug 23, 2018

Vlad Zhukov, Eugene Golikov, Maksim Kretov

Figure 1 for Differentiable lower bound for expected BLEU score

Abstract:In natural language processing tasks performance of the models is often measured with some non-differentiable metric, such as BLEU score. To use efficient gradient-based methods for optimization, it is a common workaround to optimize some surrogate loss function. This approach is effective if optimization of such loss also results in improving target metric. The corresponding problem is referred to as loss-evaluation mismatch. In the present work we propose a method for calculation of differentiable lower bound of expected BLEU score that does not involve computationally expensive sampling procedure such as the one required when using REINFORCE rule from reinforcement learning (RL) framework.

* Presented at NIPS 2017 Workshop on Conversational AI: Today's Practice and Tomorrow's Potential

Via

Access Paper or Ask Questions

Using stochastic computation graphs formalism for optimization of sequence-to-sequence model

Dec 15, 2017

Eugene Golikov, Vlad Zhukov, Maksim Kretov

Figure 1 for Using stochastic computation graphs formalism for optimization of sequence-to-sequence model

Figure 2 for Using stochastic computation graphs formalism for optimization of sequence-to-sequence model

Figure 3 for Using stochastic computation graphs formalism for optimization of sequence-to-sequence model

Abstract:Variety of machine learning problems can be formulated as an optimization task for some (surrogate) loss function. Calculation of loss function can be viewed in terms of stochastic computation graphs (SCG). We use this formalism to analyze a problem of optimization of famous sequence-to-sequence model with attention and propose reformulation of the task. Examples are given for machine translation (MT). Our work provides a unified view on different optimization approaches for sequence-to-sequence models and could help researchers in developing new network architectures with embedded stochastic nodes.

* Presented at 10th NIPS Workshop on Optimization for Machine Learning (NIPS 2017)

Via

Access Paper or Ask Questions