Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Reasoning or a Semblance of it? A Diagnostic Study of Transitive Reasoning in LLMs

Oct 26, 2024

Houman Mehrafarin, Arash Eshghi, Ioannis Konstas

Figure 1 for Reasoning or a Semblance of it? A Diagnostic Study of Transitive Reasoning in LLMs

Figure 2 for Reasoning or a Semblance of it? A Diagnostic Study of Transitive Reasoning in LLMs

Figure 3 for Reasoning or a Semblance of it? A Diagnostic Study of Transitive Reasoning in LLMs

Figure 4 for Reasoning or a Semblance of it? A Diagnostic Study of Transitive Reasoning in LLMs

Share this with someone who'll enjoy it:

Abstract:Evaluating Large Language Models (LLMs) on reasoning benchmarks demonstrates their ability to solve compositional questions. However, little is known of whether these models engage in genuine logical reasoning or simply rely on implicit cues to generate answers. In this paper, we investigate the transitive reasoning capabilities of two distinct LLM architectures, LLaMA 2 and Flan-T5, by manipulating facts within two compositional datasets: QASC and Bamboogle. We controlled for potential cues that might influence the models' performance, including (a) word/phrase overlaps across sections of test input; (b) models' inherent knowledge during pre-training or fine-tuning; and (c) Named Entities. Our findings reveal that while both models leverage (a), Flan-T5 shows more resilience to experiments (b and c), having less variance than LLaMA 2. This suggests that models may develop an understanding of transitivity through fine-tuning on knowingly relevant datasets, a hypothesis we leave to future work.

* To appear in EMNLP Main 2024

View paper on

Share this with someone who'll enjoy it:

Title:Reasoning or a Semblance of it? A Diagnostic Study of Transitive Reasoning in LLMs

Paper and Code