Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Learning from Red Teaming: Gender Bias Provocation and Mitigation in Large Language Models

Oct 17, 2023

Hsuan Su, Cheng-Chu Cheng, Hua Farn, Shachi H Kumar, Saurav Sahay, Shang-Tse Chen, Hung-yi Lee

Figure 1 for Learning from Red Teaming: Gender Bias Provocation and Mitigation in Large Language Models

Figure 2 for Learning from Red Teaming: Gender Bias Provocation and Mitigation in Large Language Models

Figure 3 for Learning from Red Teaming: Gender Bias Provocation and Mitigation in Large Language Models

Figure 4 for Learning from Red Teaming: Gender Bias Provocation and Mitigation in Large Language Models

Share this with someone who'll enjoy it:

Abstract:Recently, researchers have made considerable improvements in dialogue systems with the progress of large language models (LLMs) such as ChatGPT and GPT-4. These LLM-based chatbots encode the potential biases while retaining disparities that can harm humans during interactions. The traditional biases investigation methods often rely on human-written test cases. However, these test cases are usually expensive and limited. In this work, we propose a first-of-its-kind method that automatically generates test cases to detect LLMs' potential gender bias. We apply our method to three well-known LLMs and find that the generated test cases effectively identify the presence of biases. To address the biases identified, we propose a mitigation strategy that uses the generated test cases as demonstrations for in-context learning to circumvent the need for parameter fine-tuning. The experimental results show that LLMs generate fairer responses with the proposed approach.

View paper on

Share this with someone who'll enjoy it:

Title:Learning from Red Teaming: Gender Bias Provocation and Mitigation in Large Language Models

Paper and Code