Riaz, Adnan
(2026)
Creating a code assistance tool using low computational resources, [Dissertation thesis], Alma Mater Studiorum Università di Bologna.
Dottorato di ricerca in
Computer science and engineering, 38 Ciclo. DOI 10.48676/unibo/amsdottorato/12566.
Documenti full-text disponibili:
![Adnan_Riaz_Final_Upload.pdf [thumbnail of Adnan_Riaz_Final_Upload.pdf]](https://amsdottorato.unibo.it/style/images/fileicons/application_pdf.png) |
Documento PDF (English)
- Richiede un lettore di PDF come Xpdf o Adobe Acrobat Reader
Disponibile con Licenza: Salvo eventuali più ampie autorizzazioni dell'autore, la tesi può essere liberamente consultata e può essere effettuato il salvataggio e la stampa di una copia per fini strettamente personali di studio, di ricerca e di insegnamento, con espresso divieto di qualunque utilizzo direttamente o indirettamente commerciale. Ogni altro diritto sul materiale è riservato.
Download (2MB)
|
Abstract
This doctoral research investigates the development of a coding assistant using limited computational resources, conducted in response to a problem stated by Coesia, a leading industrial firm of automation and packaging solutions. The research addresses the need for an in-house, privacy-preserving, and resource-efficient code assistant, capable of supporting the Structured Text programming language defined by the IEC 61131-3 standard. Given the limited availability of domain-specific data and the stringent confidentiality policies typical of industrial environments, this study explores both the technical and organizational dimensions of how to build lightweight, specialized LLMs that perform effectively in industrial settings. It first presents a feasibility study of PEFTpyCODER, a Parameter-Efficient Fine-Tuned model developed as an in-house code assistant using limited computational resources. This study employs PEFT, Low-Ranked Adaptation (LoRA), and 4-bit quantization techniques for fine-tuning an open-source model. The proposed model demonstrated superior performance on two widely accepted Python benchmarks (MBPP and HumanEval) compared to SOTA code LLMs that are up to 5 times larger.
Building on this, a synthetic ST code generation pipeline is designed to create a high-quality, domain-specific dataset for PLC programming. This pipeline enables the large-scale creation and validation of a syntactically and functionally correct ST code dataset, forming the foundation for model training and evaluation. This work also examines the licensing and innovation management aspects of deploying LLMs in industrial settings, providing a broader organizational perspective. Finally, this work reports experimental results on open-source code LLMs for the ST code generation tasks using few-shot prompting, both without and with prompt optimization, and also fine-tuning of open-source code LLMs using PEFT, LoRA, and quantization techniques for ST code completion tasks. The initial results obtained from the experiments demonstrate the technical feasibility and performance of the proposed approach.
Abstract
This doctoral research investigates the development of a coding assistant using limited computational resources, conducted in response to a problem stated by Coesia, a leading industrial firm of automation and packaging solutions. The research addresses the need for an in-house, privacy-preserving, and resource-efficient code assistant, capable of supporting the Structured Text programming language defined by the IEC 61131-3 standard. Given the limited availability of domain-specific data and the stringent confidentiality policies typical of industrial environments, this study explores both the technical and organizational dimensions of how to build lightweight, specialized LLMs that perform effectively in industrial settings. It first presents a feasibility study of PEFTpyCODER, a Parameter-Efficient Fine-Tuned model developed as an in-house code assistant using limited computational resources. This study employs PEFT, Low-Ranked Adaptation (LoRA), and 4-bit quantization techniques for fine-tuning an open-source model. The proposed model demonstrated superior performance on two widely accepted Python benchmarks (MBPP and HumanEval) compared to SOTA code LLMs that are up to 5 times larger.
Building on this, a synthetic ST code generation pipeline is designed to create a high-quality, domain-specific dataset for PLC programming. This pipeline enables the large-scale creation and validation of a syntactically and functionally correct ST code dataset, forming the foundation for model training and evaluation. This work also examines the licensing and innovation management aspects of deploying LLMs in industrial settings, providing a broader organizational perspective. Finally, this work reports experimental results on open-source code LLMs for the ST code generation tasks using few-shot prompting, both without and with prompt optimization, and also fine-tuning of open-source code LLMs using PEFT, LoRA, and quantization techniques for ST code completion tasks. The initial results obtained from the experiments demonstrate the technical feasibility and performance of the proposed approach.
Tipologia del documento
Tesi di dottorato
Autore
Riaz, Adnan
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
38
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
Code Generation; Limited Computational Resources; Structured Text (ST); Code Large Language Models (Code LLMs); Parameter-Efficient Fine-Tuning (PEFT); Low-Rank Adaptation (LoRA); In-House Code
Assistant; Licensing and Innovation Management
DOI
10.48676/unibo/amsdottorato/12566
Data di discussione
26 Marzo 2026
URI
Altri metadati
Tipologia del documento
Tesi di dottorato
Autore
Riaz, Adnan
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
38
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
Code Generation; Limited Computational Resources; Structured Text (ST); Code Large Language Models (Code LLMs); Parameter-Efficient Fine-Tuning (PEFT); Low-Rank Adaptation (LoRA); In-House Code
Assistant; Licensing and Innovation Management
DOI
10.48676/unibo/amsdottorato/12566
Data di discussione
26 Marzo 2026
URI
Statistica sui download
Gestione del documento: