Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Dissecting Chain-of-Thought: A Study on Compositional In-Context Learning of MLPs

May 30, 2023

Yingcong Li, Kartik Sreenivasan, Angeliki Giannou, Dimitris Papailiopoulos, Samet Oymak

Figure 1 for Dissecting Chain-of-Thought: A Study on Compositional In-Context Learning of MLPs

Figure 2 for Dissecting Chain-of-Thought: A Study on Compositional In-Context Learning of MLPs

Figure 3 for Dissecting Chain-of-Thought: A Study on Compositional In-Context Learning of MLPs

Figure 4 for Dissecting Chain-of-Thought: A Study on Compositional In-Context Learning of MLPs

Share this with someone who'll enjoy it:

Abstract:Chain-of-thought (CoT) is a method that enables language models to handle complex reasoning tasks by decomposing them into simpler steps. Despite its success, the underlying mechanics of CoT are not yet fully understood. In an attempt to shed light on this, our study investigates the impact of CoT on the ability of transformers to in-context learn a simple to study, yet general family of compositional functions: multi-layer perceptrons (MLPs). In this setting, we reveal that the success of CoT can be attributed to breaking down in-context learning of a compositional function into two distinct phases: focusing on data related to each step of the composition and in-context learning the single-step composition function. Through both experimental and theoretical evidence, we demonstrate how CoT significantly reduces the sample complexity of in-context learning (ICL) and facilitates the learning of complex functions that non-CoT methods struggle with. Furthermore, we illustrate how transformers can transition from vanilla in-context learning to mastering a compositional function with CoT by simply incorporating an additional layer that performs the necessary filtering for CoT via the attention mechanism. In addition to these test-time benefits, we highlight how CoT accelerates pretraining by learning shortcuts to represent complex functions and how filtering plays an important role in pretraining. These findings collectively provide insights into the mechanics of CoT, inviting further investigation of its role in complex reasoning tasks.

View paper on

Share this with someone who'll enjoy it:

Title:Dissecting Chain-of-Thought: A Study on Compositional In-Context Learning of MLPs

Paper and Code