Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Stuti Mehta

Identifying Semantically Difficult Samples to Improve Text Classification

Feb 13, 2023

Shashank Mujumdar, Stuti Mehta, Hima Patel, Suman Mitra

Figure 1 for Identifying Semantically Difficult Samples to Improve Text Classification

Figure 2 for Identifying Semantically Difficult Samples to Improve Text Classification

Figure 3 for Identifying Semantically Difficult Samples to Improve Text Classification

Figure 4 for Identifying Semantically Difficult Samples to Improve Text Classification

Abstract:In this paper, we investigate the effect of addressing difficult samples from a given text dataset on the downstream text classification task. We define difficult samples as being non-obvious cases for text classification by analysing them in the semantic embedding space; specifically - (i) semantically similar samples that belong to different classes and (ii) semantically dissimilar samples that belong to the same class. We propose a penalty function to measure the overall difficulty score of every sample in the dataset. We conduct exhaustive experiments on 13 standard datasets to show a consistent improvement of up to 9% and discuss qualitative results to show effectiveness of our approach in identifying difficult samples for a text classification model.

Via

Access Paper or Ask Questions