Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Maksim Kretov

Few-shot classification in Named Entity Recognition Task

Dec 14, 2018

Alexander Fritzler, Varvara Logacheva, Maksim Kretov

Figure 1 for Few-shot classification in Named Entity Recognition Task

Figure 2 for Few-shot classification in Named Entity Recognition Task

Figure 3 for Few-shot classification in Named Entity Recognition Task

Figure 4 for Few-shot classification in Named Entity Recognition Task

Abstract:For many natural language processing (NLP) tasks the amount of annotated data is limited. This urges a need to apply semi-supervised learning techniques, such as transfer learning or meta-learning. In this work we tackle Named Entity Recognition (NER) task using Prototypical Network - a metric learning technique. It learns intermediate representations of words which cluster well into named entity classes. This property of the model allows classifying words with extremely limited number of training examples, and can potentially be used as a zero-shot learning method. By coupling this technique with transfer learning we achieve well-performing classifiers trained on only 20 instances of a target class.

* In proceedings of the 34th ACM/SIGAPP Symposium on Applied Computing

Via

Access Paper or Ask Questions

Embedding-reparameterization procedure for manifold-valued latent variables in generative models

Dec 06, 2018

Eugene Golikov, Maksim Kretov

Figure 1 for Embedding-reparameterization procedure for manifold-valued latent variables in generative models

Figure 2 for Embedding-reparameterization procedure for manifold-valued latent variables in generative models

Figure 3 for Embedding-reparameterization procedure for manifold-valued latent variables in generative models

Figure 4 for Embedding-reparameterization procedure for manifold-valued latent variables in generative models

Abstract:Conventional prior for Variational Auto-Encoder (VAE) is a Gaussian distribution. Recent works demonstrated that choice of prior distribution affects learning capacity of VAE models. We propose a general technique (embedding-reparameterization procedure, or ER) for introducing arbitrary manifold-valued variables in VAE model. We compare our technique with a conventional VAE on a toy benchmark problem. This is work in progress.

* Presented at Bayesian Deep Learning workshop (NeurIPS 2018)

Via

Access Paper or Ask Questions

Differentiable lower bound for expected BLEU score

Aug 23, 2018

Vlad Zhukov, Eugene Golikov, Maksim Kretov

Figure 1 for Differentiable lower bound for expected BLEU score

Abstract:In natural language processing tasks performance of the models is often measured with some non-differentiable metric, such as BLEU score. To use efficient gradient-based methods for optimization, it is a common workaround to optimize some surrogate loss function. This approach is effective if optimization of such loss also results in improving target metric. The corresponding problem is referred to as loss-evaluation mismatch. In the present work we propose a method for calculation of differentiable lower bound of expected BLEU score that does not involve computationally expensive sampling procedure such as the one required when using REINFORCE rule from reinforcement learning (RL) framework.

* Presented at NIPS 2017 Workshop on Conversational AI: Today's Practice and Tomorrow's Potential

Via

Access Paper or Ask Questions

Using stochastic computation graphs formalism for optimization of sequence-to-sequence model

Dec 15, 2017

Eugene Golikov, Vlad Zhukov, Maksim Kretov

Figure 1 for Using stochastic computation graphs formalism for optimization of sequence-to-sequence model

Figure 2 for Using stochastic computation graphs formalism for optimization of sequence-to-sequence model

Figure 3 for Using stochastic computation graphs formalism for optimization of sequence-to-sequence model

Abstract:Variety of machine learning problems can be formulated as an optimization task for some (surrogate) loss function. Calculation of loss function can be viewed in terms of stochastic computation graphs (SCG). We use this formalism to analyze a problem of optimization of famous sequence-to-sequence model with attention and propose reformulation of the task. Examples are given for machine translation (MT). Our work provides a unified view on different optimization approaches for sequence-to-sequence models and could help researchers in developing new network architectures with embedded stochastic nodes.

* Presented at 10th NIPS Workshop on Optimization for Machine Learning (NIPS 2017)

Via

Access Paper or Ask Questions