Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Mammad Hajili

Open foundation models for Azerbaijani language

Jul 02, 2024

Jafar Isbarov, Kavsar Huseynova, Elvin Mammadov, Mammad Hajili

Abstract:The emergence of multilingual large language models has enabled the development of language understanding and generation systems in Azerbaijani. However, most of the production-grade systems rely on cloud solutions, such as GPT-4. While there have been several attempts to develop open foundation models for Azerbaijani, these works have not found their way into common use due to a lack of systemic benchmarking. This paper encompasses several lines of work that promote open-source foundation models for Azerbaijani. We introduce (1) a large text corpus for Azerbaijani, (2) a family of encoder-only language models trained on this dataset, (3) labeled datasets for evaluating these models, and (4) extensive evaluation that covers all major open-source models with Azerbaijani support.

* Accepted to the 1st SIGTURK Workshop

Via

Access Paper or Ask Questions

A Large-Scale Study of Machine Translation in the Turkic Languages

Sep 09, 2021

Jamshidbek Mirzakhalov, Anoop Babu, Duygu Ataman, Sherzod Kariev, Francis Tyers, Otabek Abduraufov, Mammad Hajili, Sardana Ivanova, Abror Khaytbaev, Antonio Laverghetta Jr.(+6 more)

Figure 1 for A Large-Scale Study of Machine Translation in the Turkic Languages

Figure 2 for A Large-Scale Study of Machine Translation in the Turkic Languages

Figure 3 for A Large-Scale Study of Machine Translation in the Turkic Languages

Figure 4 for A Large-Scale Study of Machine Translation in the Turkic Languages

Abstract:Recent advances in neural machine translation (NMT) have pushed the quality of machine translation systems to the point where they are becoming widely adopted to build competitive systems. However, there is still a large number of languages that are yet to reap the benefits of NMT. In this paper, we provide the first large-scale case study of the practical application of MT in the Turkic language family in order to realize the gains of NMT for Turkic languages under high-resource to extremely low-resource scenarios. In addition to presenting an extensive analysis that identifies the bottlenecks towards building competitive systems to ameliorate data scarcity, our study has several key contributions, including, i) a large parallel corpus covering 22 Turkic languages consisting of common public datasets in combination with new datasets of approximately 2 million parallel sentences, ii) bilingual baselines for 26 language pairs, iii) novel high-quality test sets in three different translation domains and iv) human evaluation scores. All models, scripts, and data will be released to the public.

* 9 pages, 1 figure, 8 tables. Main proceedings of EMNLP 2021

Via

Access Paper or Ask Questions