Italiani, Paolo
(2026)
Optimization in large language models: innovative approaches leveraging knowledge distillation and pruning, [Dissertation thesis], Alma Mater Studiorum Università di Bologna.
Dottorato di ricerca in
Computer science and engineering, 38 Ciclo.
Documenti full-text disponibili:
Abstract
Large Language Models (LLMs) exhibit impressive capabilities in crucial tasks such as natural language understanding, generation, and complex reasoning, holding the potential to significantly influence our society. Nevertheless, these capabilities are accompanied by substantial computational and memory demands, underscoring the pressing need to devise effective techniques that address their efficiency challenges.
This thesis investigates a spectrum of optimization methodologies aimed at improving the efficiency and performance of LLMs. It first explores two principal paradigms, token pruning, which dynamically reduces redundant computation during inference, and knowledge distillation, which transfers knowledge from a larger teacher model to a smaller, more efficient student. Both techniques target efficiency, seeking to minimize resource consumption while maintaining competitive performance.
Beyond these compression-oriented strategies, the thesis also examines a complementary dimension of optimization: architectural refinement. This line of inquiry focuses on enhancing model performance through minimal yet strategically chosen structural modifications. Rather than compressing or accelerating models, these methods aim to optimize their representational and inductive capacities via subtle, computationally economical adjustments, such as introducing gating mechanisms, reparameterizing attention, or altering normalization strategies. The unifying principle is that small architectural interventions can yield measurable improvements in generalization, robustness, and accuracy.
Collectively, these approaches illuminate a broader understanding of LLM optimization—one that spans both efficiency-driven compression and performance-driven architectural enhancement—demonstrating that thoughtful design at multiple levels of abstraction can lead to more capable and resource-conscious language models.
Abstract
Large Language Models (LLMs) exhibit impressive capabilities in crucial tasks such as natural language understanding, generation, and complex reasoning, holding the potential to significantly influence our society. Nevertheless, these capabilities are accompanied by substantial computational and memory demands, underscoring the pressing need to devise effective techniques that address their efficiency challenges.
This thesis investigates a spectrum of optimization methodologies aimed at improving the efficiency and performance of LLMs. It first explores two principal paradigms, token pruning, which dynamically reduces redundant computation during inference, and knowledge distillation, which transfers knowledge from a larger teacher model to a smaller, more efficient student. Both techniques target efficiency, seeking to minimize resource consumption while maintaining competitive performance.
Beyond these compression-oriented strategies, the thesis also examines a complementary dimension of optimization: architectural refinement. This line of inquiry focuses on enhancing model performance through minimal yet strategically chosen structural modifications. Rather than compressing or accelerating models, these methods aim to optimize their representational and inductive capacities via subtle, computationally economical adjustments, such as introducing gating mechanisms, reparameterizing attention, or altering normalization strategies. The unifying principle is that small architectural interventions can yield measurable improvements in generalization, robustness, and accuracy.
Collectively, these approaches illuminate a broader understanding of LLM optimization—one that spans both efficiency-driven compression and performance-driven architectural enhancement—demonstrating that thoughtful design at multiple levels of abstraction can lead to more capable and resource-conscious language models.
Tipologia del documento
Tesi di dottorato
Autore
Italiani, Paolo
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
38
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
artificial intelligence, natural language processing, large language models, knowledge distillation, token pruning, model optimization, architectural refinement, efficiency, generalization, robustness.
Data di discussione
25 Marzo 2026
URI
Altri metadati
Tipologia del documento
Tesi di dottorato
Autore
Italiani, Paolo
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
38
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
artificial intelligence, natural language processing, large language models, knowledge distillation, token pruning, model optimization, architectural refinement, efficiency, generalization, robustness.
Data di discussione
25 Marzo 2026
URI
Statistica sui download
Gestione del documento: