Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Harald C. Gall

Towards Top-Down Deep Code Generation in Limited Scopes

Sep 04, 2022

Jian Gu, Harald C. Gall

Figure 1 for Towards Top-Down Deep Code Generation in Limited Scopes

Figure 2 for Towards Top-Down Deep Code Generation in Limited Scopes

Figure 3 for Towards Top-Down Deep Code Generation in Limited Scopes

Figure 4 for Towards Top-Down Deep Code Generation in Limited Scopes

Abstract:Deep code generation is a topic of deep learning for software engineering (DL4SE), which adopts neural models to generate code for the intended functions. Since end-to-end neural methods lack the awareness of domain knowledge and software hierarchy, the results often require manual correction. To systematically explore the potential improvements of code generation, we let it participate in the whole top-down development from intentions to realizations, which is possible in limited scopes. In the process, it benefits from massive samples, features, and knowledge. As the foundation, we suggest building a taxonomy on code data, namely code taxonomy, leveraging the categorization of code information. Moreover, we introduce a three-layer semantic pyramid (SP) to associate text data and code data. It identifies the information of different abstraction levels, and thus introduces the domain knowledge on development and reveals the hierarchy of software. Furthermore, we propose a semantic pyramid framework (SPF) as the approach, focusing on softwares of high modularity and low complexity. SPF divides the code generation process into stages and reserves spots for potential interactions. Eventually, we conceived application scopes for SPF.

* 5 pages, 3 figures, 2 tables, under revision

Via

Access Paper or Ask Questions

Assemble Foundation Models for Automatic Code Summarization

Jan 13, 2022

Jian Gu, Pasquale Salza, Harald C. Gall

Figure 1 for Assemble Foundation Models for Automatic Code Summarization

Figure 2 for Assemble Foundation Models for Automatic Code Summarization

Figure 3 for Assemble Foundation Models for Automatic Code Summarization

Figure 4 for Assemble Foundation Models for Automatic Code Summarization

Abstract:Automatic code summarization is beneficial to software development and maintenance since it reduces the burden of manual tasks. Currently, artificial intelligence is undergoing a paradigm shift. The foundation models pretrained on massive data and finetuned to downstream tasks surpass specially customized models. This trend inspired us to consider reusing foundation models instead of learning from scratch. Based on this, we propose a flexible and robust approach for automatic code summarization based on neural networks. We assemble available foundation models, such as CodeBERT and GPT-2, into a single model named AdaMo. Moreover, we utilize Gaussian noise as the simulation of contextual information to optimize the latent representation. Furthermore, we introduce two adaptive schemes from the perspective of knowledge transfer, namely continuous pretraining and intermediate finetuning, and design intermediate stage tasks for general sequence-to-sequence learning. Finally, we evaluate AdaMo against a benchmark dataset for code summarization, by comparing it with state-of-the-art models.

* 12 pages, 2 figures, 8 tables, accepted by SANER 2022, the camera-ready version

Via

Access Paper or Ask Questions