Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Víctor M. Darriba Bilbao

Surfing the modeling of PoS taggers in low-resource scenarios

Feb 04, 2024

Manuel Vilares Ferro, Víctor M. Darriba Bilbao, Francisco J. Ribadas-Pena, Jorge Graña Gil

Figure 1 for Surfing the modeling of PoS taggers in low-resource scenarios

Figure 2 for Surfing the modeling of PoS taggers in low-resource scenarios

Figure 3 for Surfing the modeling of PoS taggers in low-resource scenarios

Figure 4 for Surfing the modeling of PoS taggers in low-resource scenarios

Abstract:The recent trend towards the application of deep structured techniques has revealed the limits of huge models in natural language processing. This has reawakened the interest in traditional machine learning algorithms, which have proved still to be competitive in certain contexts, in particular low-resource settings. In parallel, model selection has become an essential task to boost performance at reasonable cost, even more so when we talk about processes involving domains where the training and/or computational resources are scarce. Against this backdrop, we evaluate the early estimation of learning curves as a practical mechanism for selecting the most appropriate model in scenarios characterized by the use of non-deep learners in resource-lean settings. On the basis of a formal approximation model previously evaluated under conditions of wide availability of training and validation resources, we study the reliability of such an approach in a different and much more demanding operationalenvironment. Using as case study the generation of PoS taggers for Galician, a language belonging to the Western Ibero-Romance group, the experimental results are consistent with our expectations.

* Mathematics 2022, 10(19), 3526
* 17 papes, 5 figures

Via

Access Paper or Ask Questions

Improving Large-Scale k-Nearest Neighbor Text Categorization with Label Autoencoders

Feb 03, 2024

Francisco J. Ribadas-Pena, Shuyuan Cao, Víctor M. Darriba Bilbao

Abstract:In this paper, we introduce a multi-label lazy learning approach to deal with automatic semantic indexing in large document collections in the presence of complex and structured label vocabularies with high inter-label correlation. The proposed method is an evolution of the traditional k-Nearest Neighbors algorithm which uses a large autoencoder trained to map the large label space to a reduced size latent space and to regenerate the predicted labels from this latent space. We have evaluated our proposal in a large portion of the MEDLINE biomedical document collection which uses the Medical Subject Headings (MeSH) thesaurus as a controlled vocabulary. In our experiments we propose and evaluate several document representation approaches and different label autoencoder configurations.

* Mathematics 2022, 10(16), 2867
* 22 pages, 4 figures

Via

Access Paper or Ask Questions