Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Maximilian W. Koppatz

Neural disambiguation of lemma and part of speech in morphologically rich languages

Jul 12, 2020

José María Hoya Quecedo, Maximilian W. Koppatz, Giacomo Furlan, Roman Yangarber

Figure 1 for Neural disambiguation of lemma and part of speech in morphologically rich languages

Figure 2 for Neural disambiguation of lemma and part of speech in morphologically rich languages

Figure 3 for Neural disambiguation of lemma and part of speech in morphologically rich languages

Figure 4 for Neural disambiguation of lemma and part of speech in morphologically rich languages

Abstract:We consider the problem of disambiguating the lemma and part of speech of ambiguous words in morphologically rich languages. We propose a method for disambiguating ambiguous words in context, using a large un-annotated corpus of text, and a morphological analyser -- with no manual disambiguation or data annotation. We assume that the morphological analyser produces multiple analyses for ambiguous words. The idea is to train recurrent neural networks on the output that the morphological analyser produces for unambiguous words. We present performance on POS and lemma disambiguation that reaches or surpasses the state of the art -- including supervised models -- using no manually annotated data. We evaluate the method on several morphologically rich languages.

* Proceedings of LREC-2020: the 12th Conference on Language Resources and Evaluation
* This paper contains corrigenda to a previously published paper (Hoya Quecedo et al., 2020). It corrects a mistake in the original evaluation setup, and the results reported in Section 6., in Tables 5, 6, and 7

Via

Access Paper or Ask Questions