Scarce data and complex texts: the role of LLMs in automatic text classification

Palmieri, Elena (2026) Scarce data and complex texts: the role of LLMs in automatic text classification, [Dissertation thesis], Alma Mater Studiorum Università di Bologna. Dottorato di ricerca in Computer science and engineering, 38 Ciclo.
Documenti full-text disponibili:
[thumbnail of palmieri_elena_tesi.pdf] Documento PDF (English) - Richiede un lettore di PDF come Xpdf o Adobe Acrobat Reader
Disponibile con Licenza: Creative Commons: Attribuzione - Non Commerciale 4.0 (CC BY-NC 4.0) .
Download (2MB)

Abstract

This thesis analyzes potential and limitations of LLMs for automatic text classification in domains where annotated data is scarce or unavailable. Although recent progress in NLP produced models trained on vast corpora and with great generalization capabilities, it remains unclear the extent to which these systems can complement or replace traditional supervised approaches, particularly in settings where domain knowledge, reasoning capabilities and context interpretation are essential. This research was guided by three domain questions. Firstly, to what extent can LLMs handle classification tasks that require domain knowledge that is not completely represented in training data? Secondly, how do LLMs perform in settings where the domain knowledge is explicit and well-documented compared to those where reasoning on long and unstructured text is required? Finally, is it possible to use LLMs as substitutes to human annotators generating “silver” datasets that are suitable for the training of domain-specific lightweight models? The results highlight that in domains where terminology is standard and knowledge is accessible, models such as GPT-4o and Gemini achieve high levels of accuracy. However, the performance drops when the task requires deeper reasoning or understanding of administrative or legal frameworks. This suggests that domain knowledge still represents a crucial constraint and that prompt engineering alone is not sufficient to compensate for its absence. To address data scarcity, this thesis introduces an approach based on the creation of silver datasets, automatically annotated through carefully designed prompts, and their use for training lightweight classifiers such as Logistic Regression or Support Vector Machines. In the Emilia-Romagna case study, the models trained on the silver dataset outperformed zero-shot Llama and Gemini, offering a sustainable and cost-effective solution for public administrations. Overall, this thesis presents empirical and methodological contributions to the role of LLMs in real-world classification, positioning them as assistive tools that enhance human expertise.

Abstract
Tipologia del documento
Tesi di dottorato
Autore
Palmieri, Elena
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
38
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
Large Language Models (LLMs), Unsupervised Text Classification, Data-Scarce Domains, Automatic Annotation, Zero-Shot Classification, Prompt Engineering, Public Administration, Lightweight Classifiers, Human-in-the-Loop Annotation
Data di discussione
25 Marzo 2026
URI

Altri metadati

Statistica sui download

Gestione del documento: Visualizza la tesi

^