Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Pirá: A Bilingual Portuguese-English Dataset for Question-Answering about the Ocean

Feb 04, 2022

André F. A. Paschoal, Paulo Pirozelli, Valdinei Freire, Karina V. Delgado, Sarajane M. Peres, Marcos M. José, Flávio Nakasato, André S. Oliveira, Anarosa A. F. Brandão, Anna H. R. Costa(+1 more)

Figure 1 for Pirá: A Bilingual Portuguese-English Dataset for Question-Answering about the Ocean

Figure 2 for Pirá: A Bilingual Portuguese-English Dataset for Question-Answering about the Ocean

Figure 3 for Pirá: A Bilingual Portuguese-English Dataset for Question-Answering about the Ocean

Figure 4 for Pirá: A Bilingual Portuguese-English Dataset for Question-Answering about the Ocean

Share this with someone who'll enjoy it:

Abstract:Current research in natural language processing is highly dependent on carefully produced corpora. Most existing resources focus on English; some resources focus on languages such as Chinese and French; few resources deal with more than one language. This paper presents the Pir\'a dataset, a large set of questions and answers about the ocean and the Brazilian coast both in Portuguese and English. Pir\'a is, to the best of our knowledge, the first QA dataset with supporting texts in Portuguese, and, perhaps more importantly, the first bilingual QA dataset that includes this language. The Pir\'a dataset consists of 2261 properly curated question/answer (QA) sets in both languages. The QA sets were manually created based on two corpora: abstracts related to the Brazilian coast and excerpts of United Nation reports about the ocean. The QA sets were validated in a peer-review process with the dataset contributors. We discuss some of the advantages as well as limitations of Pir\'a, as this new resource can support a set of tasks in NLP such as question-answering, information retrieval, and machine translation.

* CIKM '21: Proceedings of the 30th ACM International Conference on Information & Knowledge Management, 2021 * https://github.com/C4AI/Pira

View paper on

Share this with someone who'll enjoy it:

Title:Pirá: A Bilingual Portuguese-English Dataset for Question-Answering about the Ocean

Paper and Code