Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Improving Stack Overflow question title generation with copying enhanced CodeBERT model and bi-modal information

Sep 27, 2021

Fengji Zhang, Jacky Keung, Xiao Yu, Zhiwen Xie, Zhen Yang, Caoyuan Ma, Zhimin Zhang

Figure 1 for Improving Stack Overflow question title generation with copying enhanced CodeBERT model and bi-modal information

Figure 2 for Improving Stack Overflow question title generation with copying enhanced CodeBERT model and bi-modal information

Figure 3 for Improving Stack Overflow question title generation with copying enhanced CodeBERT model and bi-modal information

Figure 4 for Improving Stack Overflow question title generation with copying enhanced CodeBERT model and bi-modal information

Share this with someone who'll enjoy it:

Abstract:Context: Stack Overflow is very helpful for software developers who are seeking answers to programming problems. Previous studies have shown that a growing number of questions are of low-quality and thus obtain less attention from potential answerers. Gao et al. proposed a LSTM-based model (i.e., BiLSTM-CC) to automatically generate question titles from the code snippets to improve the question quality. However, only using the code snippets in question body cannot provide sufficient information for title generation, and LSTMs cannot capture the long-range dependencies between tokens. Objective: We propose CCBERT, a deep learning based novel model to enhance the performance of question title generation by making full use of the bi-modal information of the entire question body. Methods: CCBERT follows the encoder-decoder paradigm, and uses CodeBERT to encode the question body into hidden representations, a stacked Transformer decoder to generate predicted tokens, and an additional copy attention layer to refine the output distribution. Both the encoder and decoder perform the multi-head self-attention operation to better capture the long-range dependencies. We build a dataset containing more than 120,000 high-quality questions filtered from the data officially published by Stack Overflow to verify the effectiveness of the CCBERT model. Results: CCBERT achieves a better performance on the dataset, and especially outperforms BiLSTM-CC and a multi-purpose pre-trained model (BART) by 14% and 4% on average, respectively. Experiments on both code-only and low-resource datasets also show the superiority of CCBERT with less performance degradation, which are 40% and 13.5% for BiLSTM-CC, while 24% and 5% for CCBERT, respectively.

* submitted to Information and Software Technology

View paper on

Share this with someone who'll enjoy it:

Title:Improving Stack Overflow question title generation with copying enhanced CodeBERT model and bi-modal information

Paper and Code