Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Islam El Hosary

WideBot

WESSA at SemEval-2020 Task 9: Code-Mixed Sentiment Analysis using Transformers

Sep 21, 2020

Ahmed Sultan, Mahmoud Salim, Amina Gaber, Islam El Hosary

Figure 1 for WESSA at SemEval-2020 Task 9: Code-Mixed Sentiment Analysis using Transformers

Figure 2 for WESSA at SemEval-2020 Task 9: Code-Mixed Sentiment Analysis using Transformers

Figure 3 for WESSA at SemEval-2020 Task 9: Code-Mixed Sentiment Analysis using Transformers

Abstract:In this paper, we describe our system submitted for SemEval 2020 Task 9, Sentiment Analysis for Code-Mixed Social Media Text alongside other experiments. Our best performing system is a Transfer Learning-based model that fine-tunes "XLM-RoBERTa", a transformer-based multilingual masked language model, on monolingual English and Spanish data and Spanish-English code-mixed data. Our system outperforms the official task baseline by achieving a 70.1% average F1-Score on the official leaderboard using the test set. For later submissions, our system manages to achieve a 75.9% average F1-Score on the test set using CodaLab username "ahmed0sultan".

* Proceedings of SemEval-2020

Via

Access Paper or Ask Questions

WOLI at SemEval-2020 Task 12: Arabic Offensive Language Identification on Different Twitter Datasets

Sep 11, 2020

Yasser Otiefy, Ahmed Abdelmalek, Islam El Hosary

Figure 1 for WOLI at SemEval-2020 Task 12: Arabic Offensive Language Identification on Different Twitter Datasets

Figure 2 for WOLI at SemEval-2020 Task 12: Arabic Offensive Language Identification on Different Twitter Datasets

Figure 3 for WOLI at SemEval-2020 Task 12: Arabic Offensive Language Identification on Different Twitter Datasets

Figure 4 for WOLI at SemEval-2020 Task 12: Arabic Offensive Language Identification on Different Twitter Datasets

Abstract:Communicating through social platforms has become one of the principal means of personal communications and interactions. Unfortunately, healthy communication is often interfered by offensive language that can have damaging effects on the users. A key to fight offensive language on social media is the existence of an automatic offensive language detection system. This paper presents the results and the main findings of SemEval-2020, Task 12 OffensEval Sub-task A Zampieri et al. (2020), on Identifying and categorising Offensive Language in Social Media. The task was based on the Arabic OffensEval dataset Mubarak et al. (2020). In this paper, we describe the system submitted by WideBot AI Lab for the shared task which ranked 10th out of 52 participants with Macro-F1 86.9% on the golden dataset under CodaLab username "yasserotiefy". We experimented with various models and the best model is a linear SVM in which we use a combination of both character and word n-grams. We also introduced a neural network approach that enhanced the predictive ability of our system that includes CNN, highway network, Bi-LSTM, and attention layers.

Via

Access Paper or Ask Questions