Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks

Sep 29, 2023

Hao Chen, Jindong Wang, Ankit Shah, Ran Tao, Hongxin Wei, Xing Xie, Masashi Sugiyama, Bhiksha Raj

Figure 1 for Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks

Figure 2 for Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks

Figure 3 for Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks

Figure 4 for Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks

Share this with someone who'll enjoy it:

Abstract:Pre-training on large-scale datasets and then fine-tuning on downstream tasks have become a standard practice in deep learning. However, pre-training data often contain label noise that may adversely affect the generalization of the model. This paper aims to understand the nature of noise in pre-training datasets and to mitigate its impact on downstream tasks. More specifically, through extensive experiments of supervised pre-training models on synthetic noisy ImageNet-1K and YFCC15M datasets, we demonstrate that while slight noise in pre-training can benefit in-domain (ID) transfer performance, where the training and testing data share the same distribution, it always deteriorates out-of-domain (OOD) performance, where training and testing data distribution are different. We empirically verify that the reason behind is noise in pre-training shapes the feature space differently. We then propose a lightweight black-box tuning method (NMTune) to affine the feature space to mitigate the malignant effect of noise and improve generalization on both ID and OOD tasks, considering one may not be able to fully fine-tune or even access the pre-trained models. We conduct practical experiments on popular vision and language models that are pre-trained on noisy data for evaluation of our approach. Our analysis and results show the importance of this interesting and novel research direction, which we term Noisy Model Learning.

* 30 pages, 16 figures, 16 tables, preprint

View paper on

Share this with someone who'll enjoy it:

Title:Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks

Paper and Code