Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:GraphEval2000: Benchmarking and Improving Large Language Models on Graph Datasets

Jun 23, 2024

Qiming Wu, Zichen Chen, Will Corcoran, Misha Sra, Ambuj K. Singh

Figure 1 for GraphEval2000: Benchmarking and Improving Large Language Models on Graph Datasets

Figure 2 for GraphEval2000: Benchmarking and Improving Large Language Models on Graph Datasets

Figure 3 for GraphEval2000: Benchmarking and Improving Large Language Models on Graph Datasets

Figure 4 for GraphEval2000: Benchmarking and Improving Large Language Models on Graph Datasets

Share this with someone who'll enjoy it:

Abstract:Large language models (LLMs) have achieved remarkable success in natural language processing (NLP), demonstrating significant capabilities in processing and understanding text data. However, recent studies have identified limitations in LLMs' ability to reason about graph-structured data. To address this gap, we introduce GraphEval2000, the first comprehensive graph dataset, comprising 40 graph data structure problems along with 2000 test cases. Additionally, we introduce an evaluation framework based on GraphEval2000, designed to assess the graph reasoning abilities of LLMs through coding challenges. Our dataset categorizes test cases into four primary and four sub-categories, ensuring a comprehensive evaluation. We evaluate eight popular LLMs on GraphEval2000, revealing that LLMs exhibit a better understanding of directed graphs compared to undirected ones. While private LLMs consistently outperform open-source models, the performance gap is narrowing. Furthermore, to improve the usability of our evaluation framework, we propose Structured Symbolic Decomposition (SSD), an instruction-based method designed to enhance LLM performance on GraphEval2000. Results show that SSD improves the performance of GPT-3.5, GPT-4, and GPT-4o on complex graph problems, with an increase of 11.11\%, 33.37\%, and 33.37\%, respectively.

* Submitted to NeurIPs 2024 Dataset and Benchmark track, under review

View paper on

Share this with someone who'll enjoy it:

Title:GraphEval2000: Benchmarking and Improving Large Language Models on Graph Datasets

Paper and Code