Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Abdelati Hawwari

Part of speech tagging for code switched data

Nov 03, 2019

Fahad AlGhamdi, Giovanni Molina, Mona Diab, Thamar Solorio, Abdelati Hawwari, Victor Soto, Julia Hirschberg

Figure 1 for Part of speech tagging for code switched data

Figure 2 for Part of speech tagging for code switched data

Figure 3 for Part of speech tagging for code switched data

Figure 4 for Part of speech tagging for code switched data

Abstract:We address the problem of Part of Speech tagging (POS) in the context of linguistic code switching (CS). CS is the phenomenon where a speaker switches between two languages or variants of the same language within or across utterances, known as intra-sentential or inter-sentential CS, respectively. Processing CS data is especially challenging in intra-sentential data given state of the art monolingual NLP technology since such technology is geared toward the processing of one language at a time. In this paper we explore multiple strategies of applying state of the art POS taggers to CS data. We investigate the landscape in two CS language pairs, Spanish-English and Modern Standard Arabic-Arabic dialects. We compare the use of two POS taggers vs. a unified tagger trained on CS data. Our results show that applying a machine learning framework using two state of the art POS taggers achieves better performance compared to all other approaches that we investigate.

* Association for Computational Linguistics

Via

Access Paper or Ask Questions

Overview for the Second Shared Task on Language Identification in Code-Switched Data

Sep 28, 2019

Giovanni Molina, Fahad AlGhamdi, Mahmoud Ghoneim, Abdelati Hawwari, Nicolas Rey-Villamizar, Mona Diab, Thamar Solorio

Figure 1 for Overview for the Second Shared Task on Language Identification in Code-Switched Data

Figure 2 for Overview for the Second Shared Task on Language Identification in Code-Switched Data

Figure 3 for Overview for the Second Shared Task on Language Identification in Code-Switched Data

Figure 4 for Overview for the Second Shared Task on Language Identification in Code-Switched Data

Abstract:We present an overview of the second shared task on language identification in code-switched data. For the shared task, we had code-switched data from two different language pairs: Modern Standard Arabic-Dialectal Arabic (MSA-DA) and Spanish-English (SPA-ENG). We had a total of nine participating teams, with all teams submitting a system for SPA-ENG and four submitting for MSA-DA. Through evaluation, we found that once again language identification is more difficult for the language pair that is more closely related. We also found that this year's systems performed better overall than the systems from the previous shared task indicating overall progress in the state of the art for this task.

Via

Access Paper or Ask Questions

Creating a Large Multi-Layered Representational Repository of Linguistic Code Switched Arabic Data

Sep 28, 2019

Mona Diab, Mahmoud Ghoneim, Abdelati Hawwari, Fahad AlGhamdi, Nada AlMarwani, Mohamed Al-Badrashiny

Figure 1 for Creating a Large Multi-Layered Representational Repository of Linguistic Code Switched Arabic Data

Figure 2 for Creating a Large Multi-Layered Representational Repository of Linguistic Code Switched Arabic Data

Figure 3 for Creating a Large Multi-Layered Representational Repository of Linguistic Code Switched Arabic Data

Figure 4 for Creating a Large Multi-Layered Representational Repository of Linguistic Code Switched Arabic Data

Abstract:We present our effort to create a large Multi-Layered representational repository of Linguistic Code-Switched Arabic data. The process involves developing clear annotation standards and Guidelines, streamlining the annotation process, and implementing quality control measures. We used two main protocols for annotation: in-lab gold annotations and crowd sourcing annotations. We developed a web-based annotation tool to facilitate the management of the annotation process. The current version of the repository contains a total of 886,252 tokens that are tagged into one of sixteen code-switching tags. The data exhibits code switching between Modern Standard Arabic and Egyptian Dialectal Arabic representing three data genres: Tweets, commentaries, and discussion fora. The overall Inter-Annotator Agreement is 93.1%.

Via

Access Paper or Ask Questions