Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:HateCheck: Functional Tests for Hate Speech Detection Models

Dec 31, 2020

Paul Röttger, Bertram Vidgen, Dong Nguyen, Zeerak Waseem, Helen Margetts, Janet Pierrehumbert

Figure 1 for HateCheck: Functional Tests for Hate Speech Detection Models

Figure 2 for HateCheck: Functional Tests for Hate Speech Detection Models

Figure 3 for HateCheck: Functional Tests for Hate Speech Detection Models

Figure 4 for HateCheck: Functional Tests for Hate Speech Detection Models

Share this with someone who'll enjoy it:

Abstract:Detecting online hate is a difficult task that even state-of-the-art models struggle with. In previous research, hate speech detection models are typically evaluated by measuring their performance on held-out test data using metrics such as accuracy and F1 score. However, this approach makes it difficult to identify specific model weak points. It also risks overestimating generalisable model quality due to increasingly well-evidenced systematic gaps and biases in hate speech datasets. To enable more targeted diagnostic insights, we introduce HateCheck, a first suite of functional tests for hate speech detection models. We specify 29 model functionalities, the selection of which we motivate by reviewing previous research and through a series of interviews with civil society stakeholders. We craft test cases for each functionality and validate data quality through a structured annotation process. To illustrate HateCheck's utility, we test near-state-of-the-art transformer detection models as well as a popular commercial model, revealing critical model weaknesses.

View paper on

Share this with someone who'll enjoy it:

Title:HateCheck: Functional Tests for Hate Speech Detection Models

Paper and Code