Optimization in large language models: innovative approaches leveraging knowledge distillation and pruning

Italiani, Paolo (2026) Optimization in large language models: innovative approaches leveraging knowledge distillation and pruning, [Dissertation thesis], Alma Mater Studiorum Università di Bologna. Dottorato di ricerca in Computer science and engineering, 38 Ciclo.
Documenti full-text disponibili:
[thumbnail of italiani_paolo_tesi.pdf] Documento PDF (English) - Richiede un lettore di PDF come Xpdf o Adobe Acrobat Reader
Disponibile con Licenza: Creative Commons: Attribuzione 4.0 (CC BY 4.0) .
Download (23MB)

Abstract

Large Language Models (LLMs) exhibit impressive capabilities in crucial tasks such as natural language understanding, generation, and complex reasoning, holding the potential to significantly influence our society. Nevertheless, these capabilities are accompanied by substantial computational and memory demands, underscoring the pressing need to devise effective techniques that address their efficiency challenges. This thesis investigates a spectrum of optimization methodologies aimed at improving the efficiency and performance of LLMs. It first explores two principal paradigms, token pruning, which dynamically reduces redundant computation during inference, and knowledge distillation, which transfers knowledge from a larger teacher model to a smaller, more efficient student. Both techniques target efficiency, seeking to minimize resource consumption while maintaining competitive performance. Beyond these compression-oriented strategies, the thesis also examines a complementary dimension of optimization: architectural refinement. This line of inquiry focuses on enhancing model performance through minimal yet strategically chosen structural modifications. Rather than compressing or accelerating models, these methods aim to optimize their representational and inductive capacities via subtle, computationally economical adjustments, such as introducing gating mechanisms, reparameterizing attention, or altering normalization strategies. The unifying principle is that small architectural interventions can yield measurable improvements in generalization, robustness, and accuracy. Collectively, these approaches illuminate a broader understanding of LLM optimization—one that spans both efficiency-driven compression and performance-driven architectural enhancement—demonstrating that thoughtful design at multiple levels of abstraction can lead to more capable and resource-conscious language models.

Abstract
Tipologia del documento
Tesi di dottorato
Autore
Italiani, Paolo
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
38
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
artificial intelligence, natural language processing, large language models, knowledge distillation, token pruning, model optimization, architectural refinement, efficiency, generalization, robustness.
Data di discussione
25 Marzo 2026
URI

Altri metadati

Statistica sui download

Gestione del documento: Visualizza la tesi

^