Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Fine-Tuning with Divergent Chains of Thought Boosts Reasoning Through Self-Correction in Language Models

Jul 03, 2024

Haritz Puerto, Tilek Chubakov, Xiaodan Zhu, Harish Tayyar Madabushi, Iryna Gurevych

Figure 1 for Fine-Tuning with Divergent Chains of Thought Boosts Reasoning Through Self-Correction in Language Models

Figure 2 for Fine-Tuning with Divergent Chains of Thought Boosts Reasoning Through Self-Correction in Language Models

Figure 3 for Fine-Tuning with Divergent Chains of Thought Boosts Reasoning Through Self-Correction in Language Models

Figure 4 for Fine-Tuning with Divergent Chains of Thought Boosts Reasoning Through Self-Correction in Language Models

Share this with someone who'll enjoy it:

Abstract:Requiring a Large Language Model to generate intermediary reasoning steps has been shown to be an effective way of boosting performance. In fact, it has been found that instruction tuning on these intermediary reasoning steps improves model performance. In this work, we present a novel method of further improving performance by requiring models to compare multiple reasoning chains before generating a solution in a single inference step. We call this method Divergent CoT (DCoT). We find that instruction tuning on DCoT datasets boosts the performance of even smaller, and therefore more accessible, LLMs. Through a rigorous set of experiments spanning a wide range of tasks that require various reasoning types, we show that fine-tuning on DCoT consistently improves performance over the CoT baseline across model families and scales (1.3B to 70B). Through a combination of empirical and manual evaluation, we additionally show that these performance gains stem from models generating multiple divergent reasoning chains in a single inference step, indicative of the enabling of self-correction in language models. Our code and data are publicly available at https://github.com/UKPLab/arxiv2024-divergent-cot.

View paper on

Share this with someone who'll enjoy it:

Title:Fine-Tuning with Divergent Chains of Thought Boosts Reasoning Through Self-Correction in Language Models

Paper and Code