Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Mary-Joy Sidhom

Fix-A-Step: Effective Semi-supervised Learning from Uncurated Unlabeled Sets

Aug 25, 2022

Zhe Huang, Mary-Joy Sidhom, Benjamin S. Wessler, Michael C. Hughes

Figure 1 for Fix-A-Step: Effective Semi-supervised Learning from Uncurated Unlabeled Sets

Figure 2 for Fix-A-Step: Effective Semi-supervised Learning from Uncurated Unlabeled Sets

Figure 3 for Fix-A-Step: Effective Semi-supervised Learning from Uncurated Unlabeled Sets

Figure 4 for Fix-A-Step: Effective Semi-supervised Learning from Uncurated Unlabeled Sets

Abstract:Semi-supervised learning (SSL) promises gains in accuracy compared to training classifiers on small labeled datasets by also training on many unlabeled images. In realistic applications like medical imaging, unlabeled sets will be collected for expediency and thus uncurated: possibly different from the labeled set in represented classes or class frequencies. Unfortunately, modern deep SSL often makes accuracy worse when given uncurated unlabeled sets. Recent remedies suggest filtering approaches that detect out-of-distribution unlabeled examples and then discard or downweight them. Instead, we view all unlabeled examples as potentially helpful. We introduce a procedure called Fix-A-Step that can improve heldout accuracy of common deep SSL methods despite lack of curation. The key innovations are augmentations of the labeled set inspired by all unlabeled data and a modification of gradient descent updates to prevent following the multi-task SSL loss from hurting labeled-set accuracy. Though our method is simpler than alternatives, we show consistent accuracy gains on CIFAR-10 and CIFAR-100 benchmarks across all tested levels of artificial contamination for the unlabeled sets. We further suggest a real medical benchmark for SSL: recognizing the view type of ultrasound images of the heart. Our method can learn from 353,500 truly uncurated unlabeled images to deliver gains that generalize across hospitals.

Via

Access Paper or Ask Questions