Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Muhan Zeng

SWE-bench-java: A GitHub Issue Resolving Benchmark for Java

Aug 26, 2024

Daoguang Zan, Zhirong Huang, Ailun Yu, Shaoxin Lin, Yifan Shi, Wei Liu, Dong Chen, Zongshuai Qi, Hao Yu, Lei Yu(+10 more)

Figure 1 for SWE-bench-java: A GitHub Issue Resolving Benchmark for Java

Figure 2 for SWE-bench-java: A GitHub Issue Resolving Benchmark for Java

Figure 3 for SWE-bench-java: A GitHub Issue Resolving Benchmark for Java

Figure 4 for SWE-bench-java: A GitHub Issue Resolving Benchmark for Java

Abstract:GitHub issue resolving is a critical task in software engineering, recently gaining significant attention in both industry and academia. Within this task, SWE-bench has been released to evaluate issue resolving capabilities of large language models (LLMs), but has so far only focused on Python version. However, supporting more programming languages is also important, as there is a strong demand in industry. As a first step toward multilingual support, we have developed a Java version of SWE-bench, called SWE-bench-java. We have publicly released the dataset, along with the corresponding Docker-based evaluation environment and leaderboard, which will be continuously maintained and updated in the coming months. To verify the reliability of SWE-bench-java, we implement a classic method SWE-agent and test several powerful LLMs on it. As is well known, developing a high-quality multi-lingual benchmark is time-consuming and labor-intensive, so we welcome contributions through pull requests or collaboration to accelerate its iteration and refinement, paving the way for fully automated programming.

* This work is in progress

Via

Access Paper or Ask Questions

CodeR: Issue Resolving with Multi-Agent and Task Graphs

Jun 03, 2024

Dong Chen, Shaoxin Lin, Muhan Zeng, Daoguang Zan, Jian-Gang Wang, Anton Cheshkov, Jun Sun, Hao Yu, Guoliang Dong, Artem Aliev(+7 more)

Figure 1 for CodeR: Issue Resolving with Multi-Agent and Task Graphs

Figure 2 for CodeR: Issue Resolving with Multi-Agent and Task Graphs

Figure 3 for CodeR: Issue Resolving with Multi-Agent and Task Graphs

Figure 4 for CodeR: Issue Resolving with Multi-Agent and Task Graphs

Abstract:GitHub issue resolving recently has attracted significant attention from academia and industry. SWE-bench is proposed to measure the performance in resolving issues. In this paper, we propose CodeR, which adopts a multi-agent framework and pre-defined task graphs to Repair & Resolve reported bugs and add new features within code Repository. On SWE-bench lite, CodeR is able to solve 28.00% of issues, in the case of submitting only once for each issue. We examine the performance impact of each design of CodeR and offer insights to advance this research direction.

Via

Access Paper or Ask Questions

PanGu-Coder2: Boosting Large Language Models for Code with Ranking Feedback

Jul 27, 2023

Bo Shen, Jiaxin Zhang, Taihong Chen, Daoguang Zan, Bing Geng, An Fu, Muhan Zeng, Ailun Yu, Jichuan Ji, Jingyang Zhao(+2 more)

Figure 1 for PanGu-Coder2: Boosting Large Language Models for Code with Ranking Feedback

Figure 2 for PanGu-Coder2: Boosting Large Language Models for Code with Ranking Feedback

Figure 3 for PanGu-Coder2: Boosting Large Language Models for Code with Ranking Feedback

Figure 4 for PanGu-Coder2: Boosting Large Language Models for Code with Ranking Feedback

Abstract:Large Language Models for Code (Code LLM) are flourishing. New and powerful models are released on a weekly basis, demonstrating remarkable performance on the code generation task. Various approaches have been proposed to boost the code generation performance of pre-trained Code LLMs, such as supervised fine-tuning, instruction tuning, reinforcement learning, etc. In this paper, we propose a novel RRTF (Rank Responses to align Test&Teacher Feedback) framework, which can effectively and efficiently boost pre-trained large language models for code generation. Under this framework, we present PanGu-Coder2, which achieves 62.20% pass@1 on the OpenAI HumanEval benchmark. Furthermore, through an extensive evaluation on CoderEval and LeetCode benchmarks, we show that PanGu-Coder2 consistently outperforms all previous Code LLMs.

* Preprint

Via

Access Paper or Ask Questions