Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Victor Mijangos

Comparing morphological complexity of Spanish, Otomi and Nahuatl

Aug 13, 2018

Ximena Gutierrez-Vasques, Victor Mijangos

Figure 1 for Comparing morphological complexity of Spanish, Otomi and Nahuatl

Figure 2 for Comparing morphological complexity of Spanish, Otomi and Nahuatl

Figure 3 for Comparing morphological complexity of Spanish, Otomi and Nahuatl

Figure 4 for Comparing morphological complexity of Spanish, Otomi and Nahuatl

Abstract:We use two small parallel corpora for comparing the morphological complexity of Spanish, Otomi and Nahuatl. These are languages that belong to different linguistic families, the latter are low-resourced. We take into account two quantitative criteria, on one hand the distribution of types over tokens in a corpus, on the other, perplexity and entropy as indicators of word structure predictability. We show that a language can be complex in terms of how many different morphological word forms can produce, however, it may be less complex in terms of predictability of its internal structure of words.

* 7 pages, CLING 2018 Workshop on Linguistic Complexity and Natural Language Processing (LC&NLP)

Via

Access Paper or Ask Questions

Low-resource bilingual lexicon extraction using graph based word embeddings

Oct 06, 2017

Ximena Gutierrez-Vasques, Victor Mijangos

Figure 1 for Low-resource bilingual lexicon extraction using graph based word embeddings

Figure 2 for Low-resource bilingual lexicon extraction using graph based word embeddings

Figure 3 for Low-resource bilingual lexicon extraction using graph based word embeddings

Abstract:In this work we focus on the task of automatically extracting bilingual lexicon for the language pair Spanish-Nahuatl. This is a low-resource setting where only a small amount of parallel corpus is available. Most of the downstream methods do not work well under low-resources conditions. This is specially true for the approaches that use vectorial representations like Word2Vec. Our proposal is to construct bilingual word vectors from a graph. This graph is generated using translation pairs obtained from an unsupervised word alignment method. We show that, in a low-resource setting, these type of vectors are successful in representing words in a bilingual semantic space. Moreover, when a linear transformation is applied to translate words from one language to another, our graph based representations considerably outperform the popular setting that uses Word2Vec.

* Draft accepted in MICAI-IJCLA

Via

Access Paper or Ask Questions