Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Da Shen

Benchmarking Language Models for Code Syntax Understanding

Oct 26, 2022

Da Shen, Xinyun Chen, Chenguang Wang, Koushik Sen, Dawn Song

Figure 1 for Benchmarking Language Models for Code Syntax Understanding

Figure 2 for Benchmarking Language Models for Code Syntax Understanding

Figure 3 for Benchmarking Language Models for Code Syntax Understanding

Figure 4 for Benchmarking Language Models for Code Syntax Understanding

Abstract:Pre-trained language models have demonstrated impressive performance in both natural language processing and program understanding, which represent the input as a token sequence without explicitly modeling its structure. Some prior works show that pre-trained language models can capture the syntactic rules of natural languages without finetuning on syntax understanding tasks. However, there is limited understanding of how well pre-trained models understand the code structure so far. In this work, we perform the first thorough benchmarking of the state-of-the-art pre-trained models for identifying the syntactic structures of programs. Specifically, we introduce CodeSyntax, a large-scale dataset of programs annotated with the syntactic relationships in their corresponding abstract syntax trees. Our key observation is that existing language models pretrained on code still lack the understanding of code syntax. In fact, these pre-trained programming language models fail to match the performance of simple baselines based on positional offsets and keywords. We also present a natural language benchmark to highlight the differences between natural languages and programming languages in terms of syntactic structure understanding. Our findings point out key limitations of existing pre-training methods for programming languages, and suggest the importance of modeling code syntactic structures.

* Findings of EMNLP 2022

Via

Access Paper or Ask Questions

SMT-based Constraint Answer Set Solver EZSMT+

Jun 03, 2019

Da Shen, Yuliya Lierler

Figure 1 for SMT-based Constraint Answer Set Solver EZSMT+

Figure 2 for SMT-based Constraint Answer Set Solver EZSMT+

Abstract:Constraint answer set programming integrates answer set programming with constraint processing. System EZSMT+ is a constraint answer set programming tool that utilizes satisfiability modulo theory solvers for search. Its theoretical foundation lies on generalizations of Niemela's characterization of answer sets of a logic program via so called level rankings.

* This is an extended abstract submitted to LPNMR-DC 2019

Via

Access Paper or Ask Questions