Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Vijini Liyanage

An Ensemble Method Based on the Combination of Transformers with Convolutional Neural Networks to Detect Artificially Generated Text

Oct 26, 2023

Vijini Liyanage, Davide Buscaldi

Abstract:Thanks to the state-of-the-art Large Language Models (LLMs), language generation has reached outstanding levels. These models are capable of generating high quality content, thus making it a challenging task to detect generated text from human-written content. Despite the advantages provided by Natural Language Generation, the inability to distinguish automatically generated text can raise ethical concerns in terms of authenticity. Consequently, it is important to design and develop methodologies to detect artificial content. In our work, we present some classification models constructed by ensembling transformer models such as Sci-BERT, DeBERTa and XLNet, with Convolutional Neural Networks (CNNs). Our experiments demonstrate that the considered ensemble architectures surpass the performance of the individual transformer models for classification. Furthermore, the proposed SciBERT-CNN ensemble model produced an F1-score of 98.36% on the ALTA shared task 2023 data.

* In Proceedings of the 21st Annual Workshop of the Australasian Language Technology Association (ALTA 2023)

Via

Access Paper or Ask Questions

A Benchmark Corpus for the Detection of Automatically Generated Text in Academic Publications

Feb 04, 2022

Vijini Liyanage, Davide Buscaldi, Adeline Nazarenko

Figure 1 for A Benchmark Corpus for the Detection of Automatically Generated Text in Academic Publications

Figure 2 for A Benchmark Corpus for the Detection of Automatically Generated Text in Academic Publications

Figure 3 for A Benchmark Corpus for the Detection of Automatically Generated Text in Academic Publications

Figure 4 for A Benchmark Corpus for the Detection of Automatically Generated Text in Academic Publications

Abstract:Automatic text generation based on neural language models has achieved performance levels that make the generated text almost indistinguishable from those written by humans. Despite the value that text generation can have in various applications, it can also be employed for malicious tasks. The diffusion of such practices represent a threat to the quality of academic publishing. To address these problems, we propose in this paper two datasets comprised of artificially generated research content: a completely synthetic dataset and a partial text substitution dataset. In the first case, the content is completely generated by the GPT-2 model after a short prompt extracted from original papers. The partial or hybrid dataset is created by replacing several sentences of abstracts with sentences that are generated by the Arxiv-NLP model. We evaluate the quality of the datasets comparing the generated texts to aligned original texts using fluency metrics such as BLEU and ROUGE. The more natural the artificial texts seem, the more difficult they are to detect and the better is the benchmark. We also evaluate the difficulty of the task of distinguishing original from generated text by using state-of-the-art classification models.

* 9 pages including references, submitted to LREC 2022. arXiv admin note: text overlap with arXiv:2110.10577 by other authors

Via

Access Paper or Ask Questions

A Multi-language Platform for Generating Algebraic Mathematical Word Problems

Nov 19, 2019

Vijini Liyanage, Surangika Ranathunga

Figure 1 for A Multi-language Platform for Generating Algebraic Mathematical Word Problems

Figure 2 for A Multi-language Platform for Generating Algebraic Mathematical Word Problems

Figure 3 for A Multi-language Platform for Generating Algebraic Mathematical Word Problems

Figure 4 for A Multi-language Platform for Generating Algebraic Mathematical Word Problems

Abstract:Existing approaches for automatically generating mathematical word problems are deprived of customizability and creativity due to the inherent nature of template-based mechanisms they employ. We present a solution to this problem with the use of deep neural language generation mechanisms. Our approach uses a Character Level Long Short Term Memory Network (LSTM) to generate word problems, and uses POS (Part of Speech) tags to resolve the constraints found in the generated problems. Our approach is capable of generating Mathematics Word Problems in both English and Sinhala languages with an accuracy over 90%.

Via

Access Paper or Ask Questions