Documenti full-text disponibili:
Abstract
Artificial Intelligence (AI) systems are increasingly deployed in safety-critical domains, yet their internal decision-making processes often remain opaque. This thesis presents a systematic approach for analysing the internal geometry of Deep Neural Networks (DNNs) and linking it to quantifiable reliability indicators. Adopting a signal-processing perspective, DNN layers are modelled as affine operators whose learned parameters generate structured, low-rank subspaces. By projecting activations onto coordinates derived directly from these trained parameters and clustering the resulting representations, compact semantic descriptors, termed 'peepholes', are constructed. These descriptors reveal how internal evidence is organised within the network, offering interpretable summaries of its decision-making process. At the single-layer level, structural agreement measures quantify the coherence between the internal activation geometry and the model’s output decision. Extending this analysis across depth, distributed evidence maps trace the evolution of semantic support throughout the network hierarchy. A global, layer-wise structural consistency score is then derived from these maps, providing a numerical indicator of the coherence of the model’s decision process. The framework is designed to minimise external probing and perturbation. Rather than altering inputs or introducing auxiliary models, it determines whether a decision remains consistent with the structure learned by the network itself. This methodology has been validated on representative convolutional and transformer architectures and examined in the context of satellite onboard anomaly detection, where confidence assessment must coexist with strict computational budgets and operational constraints. Deviations from nominal internal structure align with reliability-related outcomes across classification errors, distributional shifts, and adversarial perturbations. Since the analysis relies solely on forward activations and parameter-induced projections, it has limited computational requirements compared to methods based on repeated perturbations, stochastic sampling, auxiliary retraining or memory-intensive reference comparisons. This makes the framework well-suited to embedded and safety-critical scenarios. This work advances trustworthy AI by grounding confidence and robustness in learned network geometry.
Abstract
Artificial Intelligence (AI) systems are increasingly deployed in safety-critical domains, yet their internal decision-making processes often remain opaque. This thesis presents a systematic approach for analysing the internal geometry of Deep Neural Networks (DNNs) and linking it to quantifiable reliability indicators. Adopting a signal-processing perspective, DNN layers are modelled as affine operators whose learned parameters generate structured, low-rank subspaces. By projecting activations onto coordinates derived directly from these trained parameters and clustering the resulting representations, compact semantic descriptors, termed 'peepholes', are constructed. These descriptors reveal how internal evidence is organised within the network, offering interpretable summaries of its decision-making process. At the single-layer level, structural agreement measures quantify the coherence between the internal activation geometry and the model’s output decision. Extending this analysis across depth, distributed evidence maps trace the evolution of semantic support throughout the network hierarchy. A global, layer-wise structural consistency score is then derived from these maps, providing a numerical indicator of the coherence of the model’s decision process. The framework is designed to minimise external probing and perturbation. Rather than altering inputs or introducing auxiliary models, it determines whether a decision remains consistent with the structure learned by the network itself. This methodology has been validated on representative convolutional and transformer architectures and examined in the context of satellite onboard anomaly detection, where confidence assessment must coexist with strict computational budgets and operational constraints. Deviations from nominal internal structure align with reliability-related outcomes across classification errors, distributional shifts, and adversarial perturbations. Since the analysis relies solely on forward activations and parameter-induced projections, it has limited computational requirements compared to methods based on repeated perturbations, stochastic sampling, auxiliary retraining or memory-intensive reference comparisons. This makes the framework well-suited to embedded and safety-critical scenarios. This work advances trustworthy AI by grounding confidence and robustness in learned network geometry.
Tipologia del documento
Tesi di dottorato
Autore
Manovi, Livia
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
38
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
Signal Processing;
Deep Neural Networks;
Explainable AI;
Trustworthy AI;
Representation Learning;
Post-hoc Interpretability;
Confidence Estimation;
Out-of-Distribution Detection;
Adversarial Robustness;
Safety-Critical Systems;
Space Systems;
Onboard Anomaly Detection
Data di discussione
10 Aprile 2026
URI
Altri metadati
Tipologia del documento
Tesi di dottorato
Autore
Manovi, Livia
Supervisore
Co-supervisore
Dottorato di ricerca
Ciclo
38
Coordinatore
Settore disciplinare
Settore concorsuale
Parole chiave
Signal Processing;
Deep Neural Networks;
Explainable AI;
Trustworthy AI;
Representation Learning;
Post-hoc Interpretability;
Confidence Estimation;
Out-of-Distribution Detection;
Adversarial Robustness;
Safety-Critical Systems;
Space Systems;
Onboard Anomaly Detection
Data di discussione
10 Aprile 2026
URI
Statistica sui download
Gestione del documento: