Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Hamzeh Motahari Khansari

HmBlogs: A big general Persian corpus

Nov 03, 2021

Hamzeh Motahari Khansari, Mehrnoush Shamsfard

Figure 1 for HmBlogs: A big general Persian corpus

Figure 2 for HmBlogs: A big general Persian corpus

Figure 3 for HmBlogs: A big general Persian corpus

Figure 4 for HmBlogs: A big general Persian corpus

Abstract:This paper introduces the hmBlogs corpus for Persian, as a low resource language. This corpus has been prepared based on a collection of nearly 20 million blog posts over a period of about 15 years from a space of Persian blogs and includes more than 6.8 billion tokens. It can be claimed that this corpus is currently the largest Persian corpus that has been prepared independently for the Persian language. This corpus is presented in both raw and preprocessed forms, and based on the preprocessed corpus some word embedding models are produced. By the provided models, the hmBlogs is compared with some of the most important corpora available in Persian, and the results show the superiority of the hmBlogs corpus over the others. These evaluations also present the importance and effects of corpora, evaluation datasets, model production methods, different hyperparameters and even the evaluation methods. In addition to evaluating the corpus and its produced language models, this research also presents a semantic analogy dataset.

* 22 pages

Via

Access Paper or Ask Questions