Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Synthetic Students: A Comparative Study of Bug Distribution Between Large Language Models and Computing Students

Oct 11, 2024

Stephen MacNeil, Magdalena Rogalska, Juho Leinonen, Paul Denny, Arto Hellas, Xandria Crosland

Figure 1 for Synthetic Students: A Comparative Study of Bug Distribution Between Large Language Models and Computing Students

Figure 2 for Synthetic Students: A Comparative Study of Bug Distribution Between Large Language Models and Computing Students

Figure 3 for Synthetic Students: A Comparative Study of Bug Distribution Between Large Language Models and Computing Students

Figure 4 for Synthetic Students: A Comparative Study of Bug Distribution Between Large Language Models and Computing Students

Share this with someone who'll enjoy it:

Abstract:Large language models (LLMs) present an exciting opportunity for generating synthetic classroom data. Such data could include code containing a typical distribution of errors, simulated student behaviour to address the cold start problem when developing education tools, and synthetic user data when access to authentic data is restricted due to privacy reasons. In this research paper, we conduct a comparative study examining the distribution of bugs generated by LLMs in contrast to those produced by computing students. Leveraging data from two previous large-scale analyses of student-generated bugs, we investigate whether LLMs can be coaxed to exhibit bug patterns that are similar to authentic student bugs when prompted to inject errors into code. The results suggest that unguided, LLMs do not generate plausible error distributions, and many of the generated errors are unlikely to be generated by real students. However, with guidance including descriptions of common errors and typical frequencies, LLMs can be shepherded to generate realistic distributions of errors in synthetic code.

View paper on

Share this with someone who'll enjoy it:

Title:Synthetic Students: A Comparative Study of Bug Distribution Between Large Language Models and Computing Students

Paper and Code