Skip to main navigation Skip to search Skip to main content

Towards automated programming feedback with open-weight language models

Research output: ThesisDoctoral ThesisCollection of Articles

Abstract

The increasing demand for computer science experts highlights the importance of supporting novice learners effectively, particularly in introductory programming courses. A key factor in maintaining learner engagement and progress is the provision of timely, constructive feedback on student code. Yet, delivering such feedback at scale remains a significant challenge: human-centered support does not scale easily, and existing automated assessment systems often lack the ability to provide meaningful and nuanced guidance. Recent advances in language model (LM) research offer promising new avenues to address this gap. Large language models (LLMs) in particular have demonstrated strong capabilities for generating educational feedback. Much of the work in this space has relied on proprietary models such as those behind ChatGPT. However, reliance on these models raises concerns around cost,control, and long-term accessibility. These challenges are motivating a shift toward open-weight alternatives. This dissertation addresses several challenges in integrating open-weight language models into educational feedback systems, presenting contributions across three complementary dimensions. First, it explores methods for enabling pre-trained LMs to repair student programs. These methods combine infilling models with search algorithms to repair programs in place and leverage automated repair tools with distillation pipelines to bootstrap training. Second, it proposes two automated evaluation approaches to assess feedback capabilities without requiring human annotation: (i) an LLM-as-a-Judge approach that leverages language models'reasoning abilities to score feedback, and (ii) a repair-as-proxy approach that uses program repair performance to measure feedback proficiency. Third, it contributes reinforcement learning techniques to align small language models (SLMs) with pedagogical objectives. These approaches allow practitioners to choose between two forms of automated supervision: alignment via model preferences (i.e., AI-generated rankings) and self-supervised alignment using verifiable program repairs. Empirical evaluations on student code datasets and public programming benchmarks demonstrate that small, open-weight models can be selected, tuned, and evaluated almost automatically to generate useful explanations, hints, and corrections for novice programmers.
Translated title of the contributionTowards automated programming feedback with open-weight language models
Original languageEnglish
QualificationDoctor's degree
Awarding Institution
  • Aalto University
Supervisors/Advisors
  • Kivelä, Mikko, Supervising Professor
  • Hellas, Arto, Thesis Advisor
  • Haaranen, Lassi, Thesis Advisor
Publisher
Print ISBNs978-952-64-2992-2
Electronic ISBNs978-952-64-2991-5
Publication statusPublished - 2026
MoE publication typeG5 Doctoral dissertation (article)

Keywords

  • language models
  • feedback
  • programming

Fingerprint

Dive into the research topics of 'Towards automated programming feedback with open-weight language models'. Together they form a unique fingerprint.

Cite this