Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Areeb Alowisheq

GLARE: Google Apps Arabic Reviews Dataset

Dec 16, 2024

Fatima AlGhamdi, Reem Mohammed, Hend Al-Khalifa, Areeb Alowisheq

Figure 1 for GLARE: Google Apps Arabic Reviews Dataset

Figure 2 for GLARE: Google Apps Arabic Reviews Dataset

Figure 3 for GLARE: Google Apps Arabic Reviews Dataset

Figure 4 for GLARE: Google Apps Arabic Reviews Dataset

Abstract:This paper introduces GLARE an Arabic Apps Reviews dataset collected from Saudi Google PlayStore. It consists of 76M reviews, 69M of which are Arabic reviews of 9,980 Android Applications. We present the data collection methodology, along with a detailed Exploratory Data Analysis (EDA) and Feature Engineering on the gathered reviews. We also highlight possible use cases and benefits of the dataset.

* Github Repo: https://github.com/Fatima-Gh/GLARE Zenodo: https://zenodo.org/records/6457824

Via

Access Paper or Ask Questions

ALLaM: Large Language Models for Arabic and English

Jul 22, 2024

M Saiful Bari, Yazeed Alnumay, Norah A. Alzahrani, Nouf M. Alotaibi, Hisham A. Alyahya, Sultan AlRashed, Faisal A. Mirza, Shaykhah Z. Alsubaie, Hassan A. Alahmed, Ghadah Alabduljabbar(+15 more)

Figure 1 for ALLaM: Large Language Models for Arabic and English

Figure 2 for ALLaM: Large Language Models for Arabic and English

Figure 3 for ALLaM: Large Language Models for Arabic and English

Figure 4 for ALLaM: Large Language Models for Arabic and English

Abstract:We present ALLaM: Arabic Large Language Model, a series of large language models to support the ecosystem of Arabic Language Technologies (ALT). ALLaM is carefully trained considering the values of language alignment and knowledge transfer at scale. Our autoregressive decoder-only architecture models demonstrate how second-language acquisition via vocabulary expansion and pretraining on a mixture of Arabic and English text can steer a model towards a new language (Arabic) without any catastrophic forgetting in the original language (English). Furthermore, we highlight the effectiveness of using parallel/translated data to aid the process of knowledge alignment between languages. Finally, we show that extensive alignment with human preferences can significantly enhance the performance of a language model compared to models of a larger scale with lower quality alignment. ALLaM achieves state-of-the-art performance in various Arabic benchmarks, including MMLU Arabic, ACVA, and Arabic Exams. Our aligned models improve both in Arabic and English from their base aligned models.

Via

Access Paper or Ask Questions

When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Feb 01, 2024

Norah Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay, Sultan Alrashed, Shaykhah Alsubaie, Yusef Almushaykeh, Faisal Mirza, Nouf Alotaibi, Nora Altwairesh, Areeb Alowisheq(+2 more)

Figure 1 for When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Figure 2 for When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Figure 3 for When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Figure 4 for When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Abstract:Large Language Model (LLM) leaderboards based on benchmark rankings are regularly used to guide practitioners in model selection. Often, the published leaderboard rankings are taken at face value - we show this is a (potentially costly) mistake. Under existing leaderboards, the relative performance of LLMs is highly sensitive to (often minute) details. We show that for popular multiple choice question benchmarks (e.g. MMLU) minor perturbations to the benchmark, such as changing the order of choices or the method of answer selection, result in changes in rankings up to 8 positions. We explain this phenomenon by conducting systematic experiments over three broad categories of benchmark perturbations and identifying the sources of this behavior. Our analysis results in several best-practice recommendations, including the advantage of a hybrid scoring method for answer selection. Our study highlights the dangers of relying on simple benchmark evaluations and charts the path for more robust evaluation schemes on the existing benchmarks.

Via

Access Paper or Ask Questions