Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:DispaRisk: Assessing and Interpreting Disparity Risks in Datasets

May 20, 2024

Jonathan Vasquez, Carlotta Domeniconi, Huzefa Rangwala

Figure 1 for DispaRisk: Assessing and Interpreting Disparity Risks in Datasets

Figure 2 for DispaRisk: Assessing and Interpreting Disparity Risks in Datasets

Figure 3 for DispaRisk: Assessing and Interpreting Disparity Risks in Datasets

Figure 4 for DispaRisk: Assessing and Interpreting Disparity Risks in Datasets

Share this with someone who'll enjoy it:

Abstract:Machine Learning algorithms (ML) impact virtually every aspect of human lives and have found use across diverse sectors, including healthcare, finance, and education. Often, ML algorithms have been found to exacerbate societal biases presented in datasets, leading to adversarial impacts on subsets/groups of individuals, in many cases minority groups. To effectively mitigate these untoward effects, it is crucial that disparities/biases are identified and assessed early in a ML pipeline. This proactive approach facilitates timely interventions to prevent bias amplification and reduce complexity at later stages of model development. In this paper, we introduce DispaRisk, a novel framework designed to proactively assess the potential risks of disparities in datasets during the initial stages of the ML pipeline. We evaluate DispaRisk's effectiveness by benchmarking it with commonly used datasets in fairness research. Our findings demonstrate the capabilities of DispaRisk to identify datasets with a high-risk of discrimination, model families prone to biases, and characteristics that heighten discrimination susceptibility in a ML pipeline. The code for our experiments is available in the following repository: https://github.com/jovasque156/disparisk

View paper on

Share this with someone who'll enjoy it:

Title:DispaRisk: Assessing and Interpreting Disparity Risks in Datasets

Paper and Code