Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Jakub Klikowski

Large Language Models in Legislative Content Analysis: A Dataset from the Polish Parliament

Mar 15, 2025

Arkadiusz Bryłkowski, Jakub Klikowski

Abstract:Large language models (LLMs) are among the best methods for processing natural language, partly due to their versatility. At the same time, domain-specific LLMs are more practical in real-life applications. This work introduces a novel natural language dataset created by acquired data from official legislative authorities' websites. The study focuses on formulating three natural language processing (NLP) tasks to evaluate the effectiveness of LLMs on legislative content analysis within the context of the Polish legal system. Key findings highlight the potential of LLMs in automating and enhancing legislative content analysis while emphasizing specific challenges, such as understanding legal context. The research contributes to the advancement of NLP in the legal field, particularly in the Polish language. It has been demonstrated that even commonly accessible data can be practically utilized for legislative content analysis.

* 15 pages, 4 figures

Via

Access Paper or Ask Questions

Employing Sentence Space Embedding for Classification of Data Stream from Fake News Domain

Jul 15, 2024

Paweł Zyblewski, Jakub Klikowski, Weronika Borek-Marciniec, Paweł Ksieniewicz

Abstract:Tabular data is considered the last unconquered castle of deep learning, yet the task of data stream classification is stated to be an equally important and demanding research area. Due to the temporal constraints, it is assumed that deep learning methods are not the optimal solution for application in this field. However, excluding the entire -- and prevalent -- group of methods seems rather rash given the progress that has been made in recent years in its development. For this reason, the following paper is the first to present an approach to natural language data stream classification using the sentence space method, which allows for encoding text into the form of a discrete digital signal. This allows the use of convolutional deep networks dedicated to image classification to solve the task of recognizing fake news based on text data. Based on the real-life Fakeddit dataset, the proposed approach was compared with state-of-the-art algorithms for data stream classification based on generalization ability and time complexity.

* 8 pages, 8 figures

Via

Access Paper or Ask Questions

WarCov -- Large multilabel and multimodal dataset from social platform

Jun 10, 2024

Weronika Borek-Marciniec, Pawel Zyblewski, Jakub Klikowski, Pawel Ksieniewicz

Figure 1 for WarCov -- Large multilabel and multimodal dataset from social platform

Figure 2 for WarCov -- Large multilabel and multimodal dataset from social platform

Figure 3 for WarCov -- Large multilabel and multimodal dataset from social platform

Figure 4 for WarCov -- Large multilabel and multimodal dataset from social platform

Abstract:In the classification tasks, from raw data acquisition to the curation of a dataset suitable for use in evaluating machine learning models, a series of steps - often associated with high costs - are necessary. In the case of Natural Language Processing, initial cleaning and conversion can be performed automatically, but obtaining labels still requires the rationalized input of human experts. As a result, even though many articles often state that "the world is filled with data", data scientists suffer from its shortage. It is crucial in the case of natural language applications, which is constantly evolving and must adapt to new concepts or events. For example, the topic of the COVID-19 pandemic and the vocabulary related to it would have been mostly unrecognizable before 2019. For this reason, creating new datasets, also in languages other than English, is still essential. This work presents a collection of 3~187~105 posts in Polish about the pandemic and the war in Ukraine published on popular social media platforms in 2022. The collection includes not only preprocessed texts but also images so it can be used also for multimodal recognition tasks. The labels define posts' topics and were created using hashtags accompanying the posts. The work presents the process of curating a dataset from acquisition to sample pattern recognition experiments.

* 13 pages, 6 figures

Via

Access Paper or Ask Questions

Hellinger Distance Weighted Ensemble for Imbalanced Data Stream Classification

Jan 30, 2021

Joanna Grzyb, Jakub Klikowski, Michał Woźniak

Figure 1 for Hellinger Distance Weighted Ensemble for Imbalanced Data Stream Classification

Figure 2 for Hellinger Distance Weighted Ensemble for Imbalanced Data Stream Classification

Figure 3 for Hellinger Distance Weighted Ensemble for Imbalanced Data Stream Classification

Figure 4 for Hellinger Distance Weighted Ensemble for Imbalanced Data Stream Classification

Abstract:The imbalanced data classification remains a vital problem. The key is to find such methods that classify both the minority and majority class correctly. The paper presents the classifier ensemble for classifying binary, non-stationary and imbalanced data streams where the Hellinger Distance is used to prune the ensemble. The paper includes an experimental evaluation of the method based on the conducted experiments. The first one checks the impact of the base classifier type on the quality of the classification. In the second experiment, the Hellinger Distance Weighted Ensemble (HDWE) method is compared to selected state-of-the-art methods using a statistical test with two base classifiers. The method was profoundly tested based on many imbalanced data streams and obtained results proved the HDWE method's usefulness.

Via

Access Paper or Ask Questions