Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Hafsteinn Einarsson

Hotter and Colder: A New Approach to Annotating Sentiment, Emotions, and Bias in Icelandic Blog Comments

Feb 24, 2025

Steinunn Rut Friðriksdóttir, Dan Saattrup Nielsen, Hafsteinn Einarsson

Abstract:This paper presents Hotter and Colder, a dataset designed to analyze various types of online behavior in Icelandic blog comments. Building on previous work, we used GPT-4o mini to annotate approximately 800,000 comments for 25 tasks, including sentiment analysis, emotion detection, hate speech, and group generalizations. Each comment was automatically labeled on a 5-point Likert scale. In a second annotation stage, comments with high or low probabilities of containing each examined behavior were subjected to manual revision. By leveraging crowdworkers to refine these automatically labeled comments, we ensure the quality and accuracy of our dataset resulting in 12,232 uniquely annotated comments and 19,301 annotations. Hotter and Colder provides an essential resource for advancing research in content moderation and automatically detectiong harmful online behaviors in Icelandic.

* To be published in the proceedings of the NoDaLiDa/Baltic-HLT 2025 conference

Via

Access Paper or Ask Questions

FoQA: A Faroese Question-Answering Dataset

Feb 11, 2025

Annika Simonsen, Dan Saattrup Nielsen, Hafsteinn Einarsson

Abstract:We present FoQA, a Faroese extractive question-answering (QA) dataset with 2,000 samples, created using a semi-automated approach combining Large Language Models (LLMs) and human validation. The dataset was generated from Faroese Wikipedia articles using GPT-4-turbo for initial QA generation, followed by question rephrasing to increase complexity and native speaker validation to ensure quality. We provide baseline performance metrics for FoQA across multiple models, including LLMs and BERT, demonstrating its effectiveness in evaluating Faroese QA performance. The dataset is released in three versions: a validated set of 2,000 samples, a complete set of all 10,001 generated samples, and a set of 2,395 rejected samples for error analysis.

* Camera-ready version for RESOURCEFUL workshop, 2025

Via

Access Paper or Ask Questions

Cross-Lingual QA as a Stepping Stone for Monolingual Open QA in Icelandic

Jul 05, 2022

Vésteinn Snæbjarnarson, Hafsteinn Einarsson

Figure 1 for Cross-Lingual QA as a Stepping Stone for Monolingual Open QA in Icelandic

Figure 2 for Cross-Lingual QA as a Stepping Stone for Monolingual Open QA in Icelandic

Figure 3 for Cross-Lingual QA as a Stepping Stone for Monolingual Open QA in Icelandic

Figure 4 for Cross-Lingual QA as a Stepping Stone for Monolingual Open QA in Icelandic

Abstract:It can be challenging to build effective open question answering (open QA) systems for languages other than English, mainly due to a lack of labeled data for training. We present a data efficient method to bootstrap such a system for languages other than English. Our approach requires only limited QA resources in the given language, along with machine-translated data, and at least a bilingual language model. To evaluate our approach, we build such a system for the Icelandic language and evaluate performance over trivia style datasets. The corpora used for training are English in origin but machine translated into Icelandic. We train a bilingual Icelandic/English language model to embed English context and Icelandic questions following methodology introduced with DensePhrases (Lee et al., 2021). The resulting system is an open domain cross-lingual QA system between Icelandic and English. Finally, the system is adapted for Icelandic only open QA, demonstrating how it is possible to efficiently create an open QA system with limited access to curated datasets in the language of interest.

Via

Access Paper or Ask Questions

Building an Icelandic Entity Linking Corpus

Jun 10, 2022

Steinunn Rut Friðriksdóttir, Valdimar Ágúst Eggertsson, Benedikt Geir Jóhannesson, Hjalti Daníelsson, Hrafn Loftsson, Hafsteinn Einarsson

Figure 1 for Building an Icelandic Entity Linking Corpus

Figure 2 for Building an Icelandic Entity Linking Corpus

Figure 3 for Building an Icelandic Entity Linking Corpus

Figure 4 for Building an Icelandic Entity Linking Corpus

Abstract:In this paper, we present the first Entity Linking corpus for Icelandic. We describe our approach of using a multilingual entity linking model (mGENRE) in combination with Wikipedia API Search (WAPIS) to label our data and compare it to an approach using WAPIS only. We find that our combined method reaches 53.9% coverage on our corpus, compared to 30.9% using only WAPIS. We analyze our results and explain the value of using a multilingual system when working with Icelandic. Additionally, we analyze the data that remain unlabeled, identify patterns and discuss why they may be more difficult to annotate.

* 9 pages, 5 figures, submitted to Dataset Creation for Lower-Resourced Languages, an LREC 2022 Workshop, 9am-1pm June 24th, 2022

Via

Access Paper or Ask Questions

A Warm Start and a Clean Crawled Corpus -- A Recipe for Good Language Models

Jan 18, 2022

Vésteinn Snæbjarnarson, Haukur Barri Símonarson, Pétur Orri Ragnarsson, Svanhvít Lilja Ingólfsdóttir, Haukur Páll Jónsson, Vilhjálmur Þorsteinsson, Hafsteinn Einarsson

Figure 1 for A Warm Start and a Clean Crawled Corpus -- A Recipe for Good Language Models

Figure 2 for A Warm Start and a Clean Crawled Corpus -- A Recipe for Good Language Models

Figure 3 for A Warm Start and a Clean Crawled Corpus -- A Recipe for Good Language Models

Figure 4 for A Warm Start and a Clean Crawled Corpus -- A Recipe for Good Language Models

Abstract:We train several language models for Icelandic, including IceBERT, that achieve state-of-the-art performance in a variety of downstream tasks, including part-of-speech tagging, named entity recognition, grammatical error detection and constituency parsing. To train the models we introduce a new corpus of Icelandic text, the Icelandic Common Crawl Corpus (IC3), a collection of high quality texts found online by targeting the Icelandic top-level-domain (TLD). Several other public data sources are also collected for a total of 16GB of Icelandic text. To enhance the evaluation of model performance and to raise the bar in baselines for Icelandic, we translate and adapt the WinoGrande dataset for co-reference resolution. Through these efforts we demonstrate that a properly cleaned crawled corpus is sufficient to achieve state-of-the-art results in NLP applications for low to medium resource languages, by comparison with models trained on a curated corpus. We further show that initializing models using existing multilingual models can lead to state-of-the-art results for some downstream tasks.

Via

Access Paper or Ask Questions

The linear hidden subset problem for the EA with scheduled and adaptive mutation rates

Aug 16, 2018

Hafsteinn Einarsson, Marcelo Matheus Gauy, Johannes Lengler, Florian Meier, Asier Mujika, Angelika Steger, Felix Weissenberger

Abstract:We study unbiased $(1+1)$ evolutionary algorithms on linear functions with an unknown number $n$ of bits with non-zero weight. Static algorithms achieve an optimal runtime of $O(n (\ln n)^{2+\epsilon})$, however, it remained unclear whether more dynamic parameter policies could yield better runtime guarantees. We consider two setups: one where the mutation rate follows a fixed schedule, and one where it may be adapted depending on the history of the run. For the first setup, we give a schedule that achieves a runtime of $(1\pm o(1))\beta n \ln n$, where $\beta \approx 3.552$, which is an asymptotic improvement over the runtime of the static setup. Moreover, we show that no schedule admits a better runtime guarantee and that the optimal schedule is essentially unique. For the second setup, we show that the runtime can be further improved to $(1\pm o(1)) e n \ln n$, which matches the performance of algorithms that know $n$ in advance. Finally, we study the related model of initial segment uncertainty with static position-dependent mutation rates, and derive asymptotically optimal lower bounds. This answers a question by Doerr, Doerr, and K\"otzing.

Via

Access Paper or Ask Questions