Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Causal Language Modeling Can Elicit Search and Reasoning Capabilities on Logic Puzzles

Sep 16, 2024

Kulin Shah, Nishanth Dikkala, Xin Wang, Rina Panigrahy

Figure 1 for Causal Language Modeling Can Elicit Search and Reasoning Capabilities on Logic Puzzles

Figure 2 for Causal Language Modeling Can Elicit Search and Reasoning Capabilities on Logic Puzzles

Figure 3 for Causal Language Modeling Can Elicit Search and Reasoning Capabilities on Logic Puzzles

Figure 4 for Causal Language Modeling Can Elicit Search and Reasoning Capabilities on Logic Puzzles

Share this with someone who'll enjoy it:

Abstract:Causal language modeling using the Transformer architecture has yielded remarkable capabilities in Large Language Models (LLMs) over the last few years. However, the extent to which fundamental search and reasoning capabilities emerged within LLMs remains a topic of ongoing debate. In this work, we study if causal language modeling can learn a complex task such as solving Sudoku puzzles. To solve a Sudoku, the model is first required to search over all empty cells of the puzzle to decide on a cell to fill and then apply an appropriate strategy to fill the decided cell. Sometimes, the application of a strategy only results in thinning down the possible values in a cell rather than concluding the exact value of the cell. In such cases, multiple strategies are applied one after the other to fill a single cell. We observe that Transformer models trained on this synthetic task can indeed learn to solve Sudokus (our model solves $94.21\%$ of the puzzles fully correctly) when trained on a logical sequence of steps taken by a solver. We find that training Transformers with the logical sequence of steps is necessary and without such training, they fail to learn Sudoku. We also extend our analysis to Zebra puzzles (known as Einstein puzzles) and show that the model solves $92.04 \%$ of the puzzles fully correctly. In addition, we study the internal representations of the trained Transformer and find that through linear probing, we can decode information about the set of possible values in any given cell from them, pointing to the presence of a strong reasoning engine implicit in the Transformer weights.

* 26 pages

View paper on

Share this with someone who'll enjoy it:

Title:Causal Language Modeling Can Elicit Search and Reasoning Capabilities on Logic Puzzles

Paper and Code