Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:ReGen: Zero-Shot Text Classification via Training Data Generation with Progressive Dense Retrieval

May 18, 2023

Yue Yu, Yuchen Zhuang, Rongzhi Zhang, Yu Meng, Jiaming Shen, Chao Zhang

Figure 1 for ReGen: Zero-Shot Text Classification via Training Data Generation with Progressive Dense Retrieval

Figure 2 for ReGen: Zero-Shot Text Classification via Training Data Generation with Progressive Dense Retrieval

Figure 3 for ReGen: Zero-Shot Text Classification via Training Data Generation with Progressive Dense Retrieval

Figure 4 for ReGen: Zero-Shot Text Classification via Training Data Generation with Progressive Dense Retrieval

Share this with someone who'll enjoy it:

Abstract:With the development of large language models (LLMs), zero-shot learning has attracted much attention for various NLP tasks. Different from prior works that generate training data with billion-scale natural language generation (NLG) models, we propose a retrieval-enhanced framework to create training data from a general-domain unlabeled corpus. To realize this, we first conduct contrastive pretraining to learn an unsupervised dense retriever for extracting the most relevant documents using class-descriptive verbalizers. We then further propose two simple strategies, namely Verbalizer Augmentation with Demonstrations and Self-consistency Guided Filtering to improve the topic coverage of the dataset while removing noisy examples. Experiments on nine datasets demonstrate that REGEN achieves 4.3% gain over the strongest baselines and saves around 70% of the time compared to baselines using large NLG models. Besides, REGEN can be naturally integrated with recently proposed large language models to boost performance.

* ACL 2023 Findings (Code: https://github.com/yueyu1030/ReGen)

View paper on

Share this with someone who'll enjoy it:

Title:ReGen: Zero-Shot Text Classification via Training Data Generation with Progressive Dense Retrieval

Paper and Code