Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Donkii: Can Annotation Error Detection Methods Find Errors in Instruction-Tuning Datasets?

Sep 04, 2023

Leon Weber-Genzel, Robert Litschko, Ekaterina Artemova, Barbara Plank

Figure 1 for Donkii: Can Annotation Error Detection Methods Find Errors in Instruction-Tuning Datasets?

Figure 2 for Donkii: Can Annotation Error Detection Methods Find Errors in Instruction-Tuning Datasets?

Figure 3 for Donkii: Can Annotation Error Detection Methods Find Errors in Instruction-Tuning Datasets?

Figure 4 for Donkii: Can Annotation Error Detection Methods Find Errors in Instruction-Tuning Datasets?

Share this with someone who'll enjoy it:

Abstract:Instruction-tuning has become an integral part of training pipelines for Large Language Models (LLMs) and has been shown to yield strong performance gains. In an orthogonal line of research, Annotation Error Detection (AED) has emerged as a tool for detecting quality issues of gold-standard labels. But so far, the application of AED methods is limited to discriminative settings. It is an open question how well AED methods generalize to generative settings which are becoming widespread via generative LLMs. In this work, we present a first and new benchmark for AED on instruction-tuning data: Donkii. It encompasses three instruction-tuning datasets enriched with annotations by experts and semi-automatic methods. We find that all three datasets contain clear-cut errors that sometimes directly propagate into instruction-tuned LLMs. We propose four AED baselines for the generative setting and evaluate them comprehensively on the newly introduced dataset. Our results demonstrate that choosing the right AED method and model size is indeed crucial, thereby deriving practical recommendations. To gain insights, we provide a first case-study to examine how the quality of the instruction-tuning datasets influences downstream performance.

View paper on

Share this with someone who'll enjoy it:

Title:Donkii: Can Annotation Error Detection Methods Find Errors in Instruction-Tuning Datasets?

Paper and Code